Vidu Q2 Reference-to-video
Vidu Q2 Reference-to-Video is an advanced AI video generation model that brings static images to life. Upload a reference image and describe the motion you want — the model generates high-quality video with smooth animation, optional audio, and cinematic quality up to 1080p.
What it does
Vidu Q2 Reference-to-video, in practice.
Vidu Q2 Reference-to-Video is a capable AI video generation model that generates video featuring specific subjects. Provide subject images alongside a motion prompt, and the model produces smooth, natural video that faithfully preserves each subject's appearance and identity — offering a strong balance of quality and cost for subject-driven workflows.
- Balanced quality and speed Solid visual consistency and motion quality at a mid-tier price point.
- Subject-driven generation Feature specific characters or objects with consistent appearance throughout the video.
- High resolution output Generate videos in 540p, 720p, or 1080p quality.
- Flexible duration Create videos from 1 to 10 seconds in length.
- Audio generation Optional audio with configurable type: full audio, speech only, or sound effects only.
- Prompt Enhancer Built-in tool to automatically improve your motion descriptions.
Run Vidu Q2 Reference-to-video
from $0.096/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
subjects | Information about the subjects in the images. Supports 1–7 subjects, total 1–7 images | |
prompt | A textual description for video generation. | default Santa Claus and the bear hug by the lakeside. |
duration | The duration of the generated media in seconds. | default 5 |
resolution | The resolution of the generated media. | 540p 720p 1080p |
audio_type | Audio type, required when audio is true, defaults to all. | all speech_only sound_effect_only |
aspect_ratio | The aspect ratio of the generated media. | 16:9 9:16 1:1 4:3 3:4 |
movement_amplitude | The movement amplitude of objects in the frame. | auto small medium large |
generate_audio | Whether to generate audio for the video. | default True |
seed | The random seed to use for the generation. -1 means a random seed will be used. | default |
Sample prompt
The prompt behind the sample.
The character rides a horse across the grassland
duration: 5resolution: 720paudio_type: allaspect_ratio: 16:9movement_amplitude: autogenerate_audio: Trueseed: FAQ
Short answers.
How much does Vidu Q2 Reference-to-video cost?
Pricing starts at $0.096 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.77. Usage is billed per request from your balance — no subscription.
Does Vidu Q2 Reference-to-video run uncensored here?
No. Vidu applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Vidu Q2 Reference-to-video for everything else it does well.
What does Vidu Q2 Reference-to-video take as input?
It is a text-to-video model. Vidu Q2 Reference-to-Video is an advanced AI video generation model that brings static images to life. Upload a reference image and describe the motion you want — the model generates high-quality video with smooth animation, optional audio, and cinematic quali
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related