Wan-2.7 Reference-to-video
Generates character-driven videos from reference images and videos, with multi-subject and voice-cloning support.
What it does
Wan-2.7 Reference-to-video, in practice.
Alibaba WAN 2.7 Reference-to-Video generates character-driven videos from reference images and videos, supporting multi-subject scenes and voice cloning.
- Character consistency: Provide reference images or videos of characters, and the model preserves their appearance across the generated video.
- Multi-subject scenes: Include up to 5 reference materials (images + videos combined) to create scenes with multiple characters interacting.
- Voice cloning: Attach a voice reference audio to transfer a character's voice into the generated video.
- Flexible framing: Five aspect ratios (16:9, 9:16, 1:1, 4:3, 3:4) at 720P or 1080P.
- Creators building character-driven stories that need consistent character identity across clips.
- Teams producing multi-character interaction videos from a set of reference assets.
Run Wan-2.7 Reference-to-video
from $0.15/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | Text prompt describing the desired video. Use labels like "character1" and "character2" to map reference materials to characters. Maximum length is 5000 characters. | |
negative_prompt | Text describing elements to exclude from the video. Maximum length is 500 characters. | |
images | Reference image URLs. Each image represents one character or subject (person, animal, object). Up to 1 image per subject. Supported formats: JPEG, JPG, PNG (transparent channel not supported), BMP, WEBP. Width and height must each be betwee | |
videos | Reference video URLs. Each video represents one character or subject and can also carry its voice. Up to 3 videos. Supported formats: mp4, mov. Duration: 1–30 seconds. Maximum file size: 100 MB per video. | |
audio | Audio URL for voice cloning. The model uses this audio as the voice for the character in the reference material. Supported formats: wav, mp3. Duration: 1–10 seconds. Maximum file size: 15 MB. Same role as `reference_voice`; `audio` is prefe | |
resolution | Output video resolution. Higher resolution increases cost. | 720P 1080P |
ratio | Aspect ratio of the generated video. | 16:9 9:16 1:1 4:3 3:4 |
duration | Video duration in seconds. Longer duration increases cost. | default 5 |
prompt_extend | Whether to use AI to enhance the prompt for better video quality. Increases generation time. | default |
seed | Random seed for video generation. Range: 0 to 2147483647. Use -1 for a random seed. | default -1 |
Sample prompt
The prompt behind the sample.
A colossal solar flare explodes beside a nearby planet, enormous waves of blazing plasma and magnetic energy spiraling through space, the planet partially illuminated by the violent stellar eruption, dramatic cosmic lighting, ultra-realistic, epic sci-fi film style.
resolution: 1080Pratio: 16:9duration: 5prompt_extend: watermark: seed: -1FAQ
Short answers.
How much does Wan-2.7 Reference-to-video cost?
Pricing starts at $0.15 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $1.20. Usage is billed per request from your balance — no subscription.
Does Wan-2.7 Reference-to-video run uncensored here?
No. Qwen applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Wan-2.7 Reference-to-video for everything else it does well.
What does Wan-2.7 Reference-to-video take as input?
It is a video-to-video model. Generates character-driven videos from reference images and videos, with multi-subject and voice-cloning support.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related