Compare · Text-to-video
MiniMax H3 Text-to-Video vs Veo3.1 Text-to-video
H3 against Veo 3.1. A large price gap, and a real difference in what each exposes as a parameter.
Every value below is read from the model catalog at generation time — price, resolution options, duration steps, aspect ratios and which parameters each model exposes. 9 of the 11 spec rows differ between these two. “Not exposed” means the model has no API parameter for it — it does not mean the capability is absent. Audio is the clearest case: MiniMax H3 generates sound with the clip but offers no switch to control it, while some models expose one.
Specs
Side by side
| MiniMax H3 Text-to-Video | Veo3.1 Text-to-video | |
|---|---|---|
| Vendor | MiniMax | |
| List price | $0.057 per second | $0.30 per second |
| Category | Text-to-video | Text-to-video |
| Resolution options | 480P, 768P, 2K | 720p, 1080p, 4k |
| Duration options | 4–15s (12 steps) | 4–8s (3 steps) |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 16:9, 9:16 |
| Audio parameter exposed | not exposed | exposed |
| Seed / reproducibility | not exposed | exposed |
| Prompt expansion | exposed | not exposed |
| Camera control parameter | not exposed | not exposed |
| Content policy | spicy route — no extra platform filter | Google vendor policy applies |
MiniMax H3 Text-to-Video is the cheaper of the two at $0.057 versus $0.30 — a 5.3× difference. At five seconds that is $0.28 against $1.50.
Output
What each one actually produces
These are the real sample clips from the catalog, with the prompt that generated each. Watch both before you decide — a spec table will not tell you whether the motion reads the way you need it to.
Sample · prompt: “Cinematic medium close-up of a desert warrior wearing a high-tech dust mask and stillsuit, glowing piercing bright blue eyes (eyes of Ibad). The wind blows fine dust across their face. Warm desert sun…”
Sample · prompt: “A car slowly driving down a quiet road, surrounded by a calm landscape, subtle motion, soft sunlight illuminating the scene, cinematic camera movement, peaceful atmosphere.”
FAQ
Short answers
Which is cheaper, MiniMax H3 Text-to-Video or Veo3.1 Text-to-video?
MiniMax H3 Text-to-Video, at $0.057 per second. The other is listed at $0.30. These are list prices; billing on this site is not live yet.
Is MiniMax H3 Text-to-Video uncensored?
Yes — this is a MiniMax H3 route served without an extra platform refusal layer. Lawful use is your responsibility and anyone under 18 is out of scope.
Is Veo3.1 Text-to-video uncensored?
No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter, see the H3 models.
What is the maximum clip length for each?
MiniMax H3 Text-to-Video: 4–15s (12 steps). Veo3.1 Text-to-video: 4–8s (3 steps). Longer work means multiple takes stitched together, not one longer request.
Which should I use for text-to-video?
Whichever exposes the parameter you need — the table above is the honest answer. If you are iterating, use the cheaper model until the motion is right, then re-render once. If you need a specific resolution path, check the resolution row: some tiers reach it natively and others via a super-resolution pass, which is not the same thing.