minimaxiH3 Get access
QwenText-to-videoVendor content policy applies

Wan-3.0 Text-to-video

All-in-one Wan3.0 renderer: cinematic, hyper-real video from a text prompt, up to 30s with smart-duration and adaptive aspect ratio.

$0.060per secondStarting price at the base resolution and quality tier.
$0.48Typical 8-second clip at the base tier. Longer or higher-resolution outputs scale with duration and tier.
Pay per useBilled per request from your balance. No subscription, no minimum.
Sample output · prompt: “The race car sped along the track.”

What it does

Wan-3.0 Text-to-video, in practice.

Wan 3.0 is Alibaba's all-in-one "super-real world renderer" — a single model that turns a text prompt into a cinematic, hyper-realistic clip with native audio. It renders not just motion but the texture of a scene, fine human detail, and even structured on-screen information (UI, text, animation) with pixel-level fidelity.

  • Native long-form generation Produce clips up to 30 seconds in a single pass — richer narrative, more creative freedom.
  • Smart duration Let the model pick the ideal length from your intent and content (`duration = -1`).
  • Hyper-real rendering Film-grade texture, lighting, and lifelike movement from a plain-text description.
  • Sound on by default The output includes a synchronized audio track; toggle it off at no extra cost.
  • Flexible aspect ratios adaptive, 16:9, 4:3, 1:1, 3:4, or 9:16.

Run Wan-3.0 Text-to-video

from $0.060/sec
resolution
ratio
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
promptThe text prompt describing the scene, subject, and action (up to 20000 characters).
resolutionOutput resolution. Native tiers: 480p, 720p, 1080p. ESR tiers: 720p-esr, 1080p-esr, 1440p-esr, 4k-esr.1080p 720p 480p 720p-esr 1080p-esr 1440p-esr 4k-esr
durationVideo length in seconds (2-30). Pass -1 for smart-duration (the model picks the best length).-1 2 3 4 5 6 7 8 9 10 11 12
ratioThe aspect ratio of the generated video. Use 'adaptive' to let the model choose.adaptive 16:9 4:3 1:1 3:4 9:16
audioWhether the output video includes an audio track. Same price either way.default True
seedThe random seed to use for the generation. -1 means a random seed will be used.

Sample prompt

The prompt behind the sample.

The race car sped along the track.
resolution: 1080Pduration: 5ratio: 16:9audio: True

FAQ

Short answers.

How much does Wan-3.0 Text-to-video cost?

Pricing starts at $0.060 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.48. Usage is billed per request from your balance — no subscription.

Does Wan-3.0 Text-to-video run uncensored here?

No. Qwen applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Wan-3.0 Text-to-video for everything else it does well.

What does Wan-3.0 Text-to-video take as input?

It is a text-to-video model. All-in-one Wan3.0 renderer: cinematic, hyper-real video from a text prompt, up to 30s with smart-duration and adaptive aspect ratio.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run Wan-3.0 Text-to-video.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models