minimaxiH3 Get access
QwenVideo-to-videoVendor content policy applies

Wan-2.7 Reference-to-video

Generates character-driven videos from reference images and videos, with multi-subject and voice-cloning support.

$0.15per secondStarting price at the base resolution and quality tier.
$1.20Typical 8-second clip at the base tier. Longer or higher-resolution outputs scale with duration and tier.
Pay per useBilled per request from your balance. No subscription, no minimum.
Sample output · prompt: “A colossal solar flare explodes beside a nearby planet, enormous waves of blazing plasma and magnetic energy spiraling through space, the planet partially illuminated by the violent stellar eruption, dramatic cosmic ligh…”

What it does

Wan-2.7 Reference-to-video, in practice.

Alibaba WAN 2.7 Reference-to-Video generates character-driven videos from reference images and videos, supporting multi-subject scenes and voice cloning.

  • Character consistency: Provide reference images or videos of characters, and the model preserves their appearance across the generated video.
  • Multi-subject scenes: Include up to 5 reference materials (images + videos combined) to create scenes with multiple characters interacting.
  • Voice cloning: Attach a voice reference audio to transfer a character's voice into the generated video.
  • Flexible framing: Five aspect ratios (16:9, 9:16, 1:1, 4:3, 3:4) at 720P or 1080P.
  • Creators building character-driven stories that need consistent character identity across clips.
  • Teams producing multi-character interaction videos from a set of reference assets.

Run Wan-2.7 Reference-to-video

from $0.15/sec
Drop a video hereor click to choose a fileFiles stay in your project. Nothing is trained on.
resolution
ratio
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
promptText prompt describing the desired video. Use labels like "character1" and "character2" to map reference materials to characters. Maximum length is 5000 characters.
negative_promptText describing elements to exclude from the video. Maximum length is 500 characters.
imagesReference image URLs. Each image represents one character or subject (person, animal, object). Up to 1 image per subject. Supported formats: JPEG, JPG, PNG (transparent channel not supported), BMP, WEBP. Width and height must each be betwee
videosReference video URLs. Each video represents one character or subject and can also carry its voice. Up to 3 videos. Supported formats: mp4, mov. Duration: 1–30 seconds. Maximum file size: 100 MB per video.
audioAudio URL for voice cloning. The model uses this audio as the voice for the character in the reference material. Supported formats: wav, mp3. Duration: 1–10 seconds. Maximum file size: 15 MB. Same role as `reference_voice`; `audio` is prefe
resolutionOutput video resolution. Higher resolution increases cost.720P 1080P
ratioAspect ratio of the generated video.16:9 9:16 1:1 4:3 3:4
durationVideo duration in seconds. Longer duration increases cost.default 5
prompt_extendWhether to use AI to enhance the prompt for better video quality. Increases generation time.default
seedRandom seed for video generation. Range: 0 to 2147483647. Use -1 for a random seed.default -1

Sample prompt

The prompt behind the sample.

A colossal solar flare explodes beside a nearby planet, enormous waves of blazing plasma and magnetic energy spiraling through space, the planet partially illuminated by the violent stellar eruption, dramatic cosmic lighting, ultra-realistic, epic sci-fi film style.
resolution: 1080Pratio: 16:9duration: 5prompt_extend: watermark: seed: -1

FAQ

Short answers.

How much does Wan-2.7 Reference-to-video cost?

Pricing starts at $0.15 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $1.20. Usage is billed per request from your balance — no subscription.

Does Wan-2.7 Reference-to-video run uncensored here?

No. Qwen applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Wan-2.7 Reference-to-video for everything else it does well.

What does Wan-2.7 Reference-to-video take as input?

It is a video-to-video model. Generates character-driven videos from reference images and videos, with multi-subject and voice-cloning support.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run Wan-2.7 Reference-to-video.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models