Wan-3.0-Prime Reference-to-video
All-in-One Reference: keep subjects consistent from any mix of reference images, videos, and audio; pixel-level identity/voice/space alignment.
What it does
Wan-3.0-Prime Reference-to-video, in practice.
Wan 3.0 Prime Reference-to-Video is the "all-in-one reference" mode: give it any mix of reference images, videos, and audio, and it keeps those subjects, props, voices, and spatial relationships consistent throughout a new clip driven by your prompt. This is pixel-level identity preservation — not "close enough," but faithful replication of the reference details.
- Mix four modalities Combine reference images, reference videos, and reference audio in one request.
- Pixel-level consistency Characters, props, voices, and spatial relationships stay aligned across the whole clip.
- Multi-subject scenes Reference several subjects at once and direct them with your prompt.
- Native long-form + audio Up to 30 seconds with a synchronized audio track.
- Flexible aspect ratios adaptive, 16:9, 4:3, 1:1, 3:4, or 9:16.
Run Wan-3.0-Prime Reference-to-video
from $0.091/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | The text prompt describing the scene and action for the referenced subjects (up to 20000 characters). | |
refers | All-in-One Reference materials. Any mix of reference images (<=10), reference videos (<=5, total <=15s), and reference audio (<=5, total <=15s), each a public URL. 'type' is optional — inferred from the URL when omitted. Image: jpeg/jpg/png | |
resolution | Output resolution. Native tiers: 480p, 720p, 1080p. ESR tiers: 720p-esr, 1080p-esr, 1440p-esr, 4k-esr. | 1080p 720p 480p 720p-esr 1080p-esr 1440p-esr 4k-esr |
duration | Video length in seconds (2-30). Pass -1 for smart-duration (the model picks the best length). | -1 2 3 4 5 6 7 8 9 10 11 12 |
ratio | The aspect ratio of the generated video. Use 'adaptive' to let the model choose from the references. | adaptive 16:9 4:3 1:1 3:4 9:16 |
audio | Whether the output video includes an audio track. Same price either way. | default True |
enable_thinking | Enable deep thinking mode. Required when providing a file or link input (document / webpage parsing). Not recommended when no file/link is provided. | default True |
file | Optional document input to parse (requires enable_thinking). Public URL. Formats: docx, doc, xlsx, xls, pptx, ppt, pdf, txt, key, pages, numbers, md; <= 100MB, <= 50 pages. Mutually exclusive with 'link'. | |
link | Optional public webpage URL to parse (requires enable_thinking). Only pages that need no login. Mutually exclusive with 'file'. | |
seed | The random seed to use for the generation. -1 means a random seed will be used. |
Sample prompt
The prompt behind the sample.
Transform the visual style into a cartoon or anime style.
resolution: 720pduration: 5ratio: 9:16enable_thinking: Trueaudio: TrueFAQ
Short answers.
How much does Wan-3.0-Prime Reference-to-video cost?
Pricing starts at $0.091 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.73. Usage is billed per request from your balance — no subscription.
Does Wan-3.0-Prime Reference-to-video run uncensored here?
No. Qwen applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Wan-3.0-Prime Reference-to-video for everything else it does well.
What does Wan-3.0-Prime Reference-to-video take as input?
It is a video-to-video model. All-in-One Reference: keep subjects consistent from any mix of reference images, videos, and audio; pixel-level identity/voice/space alignment.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related