Wan-3.0 Text-to-video
All-in-one Wan3.0 renderer: cinematic, hyper-real video from a text prompt, up to 30s with smart-duration and adaptive aspect ratio.
What it does
Wan-3.0 Text-to-video, in practice.
Wan 3.0 is Alibaba's all-in-one "super-real world renderer" — a single model that turns a text prompt into a cinematic, hyper-realistic clip with native audio. It renders not just motion but the texture of a scene, fine human detail, and even structured on-screen information (UI, text, animation) with pixel-level fidelity.
- Native long-form generation Produce clips up to 30 seconds in a single pass — richer narrative, more creative freedom.
- Smart duration Let the model pick the ideal length from your intent and content (`duration = -1`).
- Hyper-real rendering Film-grade texture, lighting, and lifelike movement from a plain-text description.
- Sound on by default The output includes a synchronized audio track; toggle it off at no extra cost.
- Flexible aspect ratios adaptive, 16:9, 4:3, 1:1, 3:4, or 9:16.
Run Wan-3.0 Text-to-video
from $0.060/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | The text prompt describing the scene, subject, and action (up to 20000 characters). | |
resolution | Output resolution. Native tiers: 480p, 720p, 1080p. ESR tiers: 720p-esr, 1080p-esr, 1440p-esr, 4k-esr. | 1080p 720p 480p 720p-esr 1080p-esr 1440p-esr 4k-esr |
duration | Video length in seconds (2-30). Pass -1 for smart-duration (the model picks the best length). | -1 2 3 4 5 6 7 8 9 10 11 12 |
ratio | The aspect ratio of the generated video. Use 'adaptive' to let the model choose. | adaptive 16:9 4:3 1:1 3:4 9:16 |
audio | Whether the output video includes an audio track. Same price either way. | default True |
seed | The random seed to use for the generation. -1 means a random seed will be used. |
Sample prompt
The prompt behind the sample.
The race car sped along the track.
resolution: 1080Pduration: 5ratio: 16:9audio: TrueFAQ
Short answers.
How much does Wan-3.0 Text-to-video cost?
Pricing starts at $0.060 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.48. Usage is billed per request from your balance — no subscription.
Does Wan-3.0 Text-to-video run uncensored here?
No. Qwen applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Wan-3.0 Text-to-video for everything else it does well.
What does Wan-3.0 Text-to-video take as input?
It is a text-to-video model. All-in-one Wan3.0 renderer: cinematic, hyper-real video from a text prompt, up to 30s with smart-duration and adaptive aspect ratio.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related