minimaxiH3 Get access
MiniMaxImage-to-videoSpicy route · no extra filter

MiniMax H3 Reference-to-Video

MiniMax H3 reference-to-video: generate a video that keeps the subject from a reference image, driven by a text prompt. Supports 2K, 5-15s.

$0.057per secondStarting price at the base resolution and quality tier.
$0.46Typical 8-second clip at the base tier. Longer or higher-resolution outputs scale with duration and tier.
Pay per useBilled per request from your balance. No subscription, no minimum.
Sample output · prompt: “Little Bear straightened his clothes, getting ready to look cool.”

What it does

MiniMax H3 Reference-to-Video, in practice.

MiniMax H3 Reference-to-Video generates a video that keeps the subjects from your reference materials consistent throughout, driven by your text prompt. Instead of animating a fixed frame, it uses the references as identity anchors — ideal for placing a character, product, or style into new scenes and motions. References can be any mix of images, videos, and an audio track (e.g. a music beat) to sync the motion to.

  • Subject consistency Preserve one or several people, characters, or objects' identities across the whole clip.
  • Mix reference types Combine reference images, reference videos, and reference audio in a single request.
  • Audio sync (optional) Provide a reference audio track and sync the action to the beat.
  • Prompt-driven scenes Put the reference subjects into any scene or action you describe.
  • High resolution output Generate videos in 2K quality.
  • Flexible aspect ratios adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16.

Run MiniMax H3 Reference-to-Video

from $0.057/sec
Drop an image hereor click to choose a fileFiles stay in your project. Nothing is trained on.
resolution
duration
ratio
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
promptThe text prompt describing the video to generate.
refersReference materials that anchor the generation. Any mix of reference images, videos, and audio, each as a public URL. At least one image OR video is required (audio alone is not allowed). Image formats: png, jpeg, jpg, webp; video: mp4, mov
resolutionThe resolution of the generated video.480P 768P 2K
durationThe duration of the generated video in seconds.4 5 6 7 8 9 10 11 12 13 14 15
ratioThe aspect ratio of the generated video. Use 'adaptive' to let the model choose.adaptive 21:9 16:9 4:3 1:1 3:4 9:16
prompt_expansionWhether to expand the prompt for better results.default

Sample prompt

The prompt behind the sample.

Little Bear straightened his clothes, getting ready to look cool.
resolution: 768Pduration: 8ratio: adaptive

FAQ

Short answers.

How much does MiniMax H3 Reference-to-Video cost?

Pricing starts at $0.057 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.46. Usage is billed per request from your balance — no subscription.

Is this the uncensored route?

Yes. This is the MiniMax H3 route served here without an extra platform refusal layer on top of the model. Lawful prompts and outputs are your responsibility; anyone under 18 is out of scope.

What does MiniMax H3 Reference-to-Video take as input?

It is a image-to-video model. MiniMax H3 reference-to-video: generate a video that keeps the subject from a reference image, driven by a text prompt. Supports 2K, 5-15s.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run MiniMax H3 Reference-to-Video.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models