MiniMax H3 Max Text-to-Video
MiniMax H3 Max text-to-video: generate a cinematic video from a text prompt. Supports 480P、768P, 5-15s., and 16:9/9:16/1:1/adaptive aspect ratios.

What it does
MiniMax H3 Max Text-to-Video, in practice.
MiniMax H3 Max Text-to-Video is a speed-optimized AI video generation model that turns a text prompt into a complete 24fps clip with synchronized audio. A 5-second 768P video renders in under 3 seconds, and even a full 15-second clip finishes in roughly 15 seconds — fast enough to iterate on an idea in real time rather than waiting on a queue.
- Near-instant generation A 5s 768P clip in under 3 seconds; 15s clips in about 15 seconds.
- Audio included Every clip is generated as complete audio-video at 24fps — no separate scoring or sound pass.
- Cinematic motion Fluid camera work and lifelike movement from a plain-text description.
- Flexible aspect ratios 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16 to fit any platform.
- Selectable duration Produce 5–15s clips at 480P or 768P.
Run MiniMax H3 Max Text-to-Video
from $0.072/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | The text prompt describing the video to generate. | |
resolution | Native 480P or 768P output, or SR-enhanced 1440P or 4K short edge. | 480P 768P 1440p-sr 4k-sr |
duration | The duration of the generated video in seconds. | 5 6 7 8 9 10 11 12 13 14 15 |
ratio | The aspect ratio of the generated video. | 21:9 16:9 4:3 1:1 3:4 9:16 |
prompt_expansion | Whether to expand the prompt for better results. | default |
Sample prompt
The prompt behind the sample.
A tiny hamster wearing a miniature explorer's outfit walks through a vast golden wheat field at sunset, carrying a small red gift box. A gentle breeze moves the wheat, warm sunlight shines through the clouds, and the hamster suddenly stops as a flock of birds flies across the glowing sky. The camera slowly tracks behind the hamster, then moves into a cinematic close-up as it looks toward a distant cozy cottage. Emotional storytelling, cinematic composition, warm golden-hour lighting, shallow depth of field, subtle film grain, realistic fur details, natural movement, atmospheric, magical yet be
resolution: 768Pduration: 8ratio: 16:9prompt_expansion: FAQ
Short answers.
How much does MiniMax H3 Max Text-to-Video cost?
Pricing starts at $0.072 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.58. Usage is billed per request from your balance — no subscription.
Is this the uncensored route?
Yes. This is the MiniMax H3 route served here without an extra platform refusal layer on top of the model. Lawful prompts and outputs are your responsibility; anyone under 18 is out of scope.
What does MiniMax H3 Max Text-to-Video take as input?
It is a text-to-video model. MiniMax H3 Max text-to-video: generate a cinematic video from a text prompt. Supports 480P、768P, 5-15s., and 16:9/9:16/1:1/adaptive aspect ratios.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related