minimaxiH3 Get access

Compare · Text-to-video

MiniMax H3 Text-to-Video vs Veo3.1 Text-to-video

H3 against Veo 3.1. A large price gap, and a real difference in what each exposes as a parameter.

Every value below is read from the model catalog at generation time — price, resolution options, duration steps, aspect ratios and which parameters each model exposes. 9 of the 11 spec rows differ between these two. “Not exposed” means the model has no API parameter for it — it does not mean the capability is absent. Audio is the clearest case: MiniMax H3 generates sound with the clip but offers no switch to control it, while some models expose one.

Specs

Side by side

9 of 11 rows differ
MiniMax H3 Text-to-VideoVeo3.1 Text-to-video
VendorMiniMaxGoogle
List price$0.057 per second$0.30 per second
CategoryText-to-videoText-to-video
Resolution options480P, 768P, 2K720p, 1080p, 4k
Duration options4–15s (12 steps)4–8s (3 steps)
Aspect ratios21:9, 16:9, 4:3, 1:1, 3:4, 9:1616:9, 9:16
Audio parameter exposednot exposedexposed
Seed / reproducibilitynot exposedexposed
Prompt expansionexposednot exposed
Camera control parameternot exposednot exposed
Content policyspicy route — no extra platform filterGoogle vendor policy applies

MiniMax H3 Text-to-Video is the cheaper of the two at $0.057 versus $0.30 — a 5.3× difference. At five seconds that is $0.28 against $1.50.

Output

What each one actually produces

These are the real sample clips from the catalog, with the prompt that generated each. Watch both before you decide — a spec table will not tell you whether the motion reads the way you need it to.

MiniMax H3 Text-to-Video · 8s
Sample · prompt: “Cinematic medium close-up of a desert warrior wearing a high-tech dust mask and stillsuit, glowing piercing bright blue eyes (eyes of Ibad). The wind blows fine dust across their face. Warm desert sun…”
Veo3.1 Text-to-video · 8s
Sample · prompt: “A car slowly driving down a quiet road, surrounded by a calm landscape, subtle motion, soft sunlight illuminating the scene, cinematic camera movement, peaceful atmosphere.”

FAQ

Short answers

Which is cheaper, MiniMax H3 Text-to-Video or Veo3.1 Text-to-video?

MiniMax H3 Text-to-Video, at $0.057 per second. The other is listed at $0.30. These are list prices; billing on this site is not live yet.

Is MiniMax H3 Text-to-Video uncensored?

Yes — this is a MiniMax H3 route served without an extra platform refusal layer. Lawful use is your responsibility and anyone under 18 is out of scope.

Is Veo3.1 Text-to-video uncensored?

No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter, see the H3 models.

What is the maximum clip length for each?

MiniMax H3 Text-to-Video: 4–15s (12 steps). Veo3.1 Text-to-video: 4–8s (3 steps). Longer work means multiple takes stitched together, not one longer request.

Which should I use for text-to-video?

Whichever exposes the parameter you need — the table above is the honest answer. If you are iterating, use the cheaper model until the motion is right, then re-render once. If you need a specific resolution path, check the resolution row: some tiers reach it natively and others via a super-resolution pass, which is not the same thing.

Run either one in Chat.

An invite code opens Chat with both models loaded. No code? Join the waitlist and name the model you want.

Get access Browse models