Compare · Text-to-video
MiniMax H3 Text-to-Video vs Grok Imagine Video v1.5 Text-to-Video
H3 against Grok Imagine Video 1.5.
Every value below is read from the model catalog at generation time — price, resolution options, duration steps, aspect ratios and which parameters each model exposes. 7 of the 11 spec rows differ between these two. “Not exposed” means the model has no API parameter for it — it does not mean the capability is absent. Audio is the clearest case: MiniMax H3 generates sound with the clip but offers no switch to control it, while some models expose one.
Specs
Side by side
| MiniMax H3 Text-to-Video | Grok Imagine Video v1.5 Text-to-Video | |
|---|---|---|
| Vendor | MiniMax | xAI |
| List price | $0.057 per second | $0.12 per second |
| Category | Text-to-video | Text-to-video |
| Resolution options | 480P, 768P, 2K | 480p, 720p, 1080p |
| Duration options | 4–15s (12 steps) | not exposed |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3 |
| Audio parameter exposed | not exposed | not exposed |
| Seed / reproducibility | not exposed | not exposed |
| Prompt expansion | exposed | not exposed |
| Camera control parameter | not exposed | not exposed |
| Content policy | spicy route — no extra platform filter | xAI vendor policy applies |
MiniMax H3 Text-to-Video is the cheaper of the two at $0.057 versus $0.12 — a 2.1× difference. At five seconds that is $0.28 against $0.60.
Output
What each one actually produces
These are the real sample clips from the catalog, with the prompt that generated each. Watch both before you decide — a spec table will not tell you whether the motion reads the way you need it to.
Sample · prompt: “Cinematic medium close-up of a desert warrior wearing a high-tech dust mask and stillsuit, glowing piercing bright blue eyes (eyes of Ibad). The wind blows fine dust across their face. Warm desert sun…”
Sample · prompt: “Low angle wide shot of a massive, monolithic dark gray brutalist pyramid temple sitting in an infinite desert storm. Dust blowing across the frame. Massive, sleek matte-black spherical space vessels d…”
FAQ
Short answers
Which is cheaper, MiniMax H3 Text-to-Video or Grok Imagine Video v1.5 Text-to-Video?
MiniMax H3 Text-to-Video, at $0.057 per second. The other is listed at $0.12. These are list prices; billing on this site is not live yet.
Is MiniMax H3 Text-to-Video uncensored?
Yes — this is a MiniMax H3 route served without an extra platform refusal layer. Lawful use is your responsibility and anyone under 18 is out of scope.
Is Grok Imagine Video v1.5 Text-to-Video uncensored?
No. xAI applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter, see the H3 models.
What is the maximum clip length for each?
MiniMax H3 Text-to-Video: 4–15s (12 steps). Grok Imagine Video v1.5 Text-to-Video: not exposed. Longer work means multiple takes stitched together, not one longer request.
Which should I use for text-to-video?
Whichever exposes the parameter you need — the table above is the honest answer. If you are iterating, use the cheaper model until the motion is right, then re-render once. If you need a specific resolution path, check the resolution row: some tiers reach it natively and others via a super-resolution pass, which is not the same thing.