Vidu Q3-Mix Reference to Video
Vidu Q3-Mix Reference-to-Video generates videos from 1-4 reference images with consistent subjects. Offers strong visual quality with intelligent scene transitions, smooth dynamic effects, and audio support up to 1080p.
What it does
Vidu Q3-Mix Reference to Video, in practice.
Vidu Q3 is an advanced AI video generation model developed by Shengshu Technology (生数科技) in collaboration with Tsinghua University. Released on January 30, 2026, Vidu Q3 is designed to produce high-fidelity, synchronized audio-visual content with industry-leading continuous video length and native support for integrated audio generation. The model represents a significant advancement in automated video synthesis by unifying multiple complex video generation tasks—such as lip-synced dialogue, dynamic camera movements, and multi-shot storytelling—into a single-pass framework. Leveraging a novel Transformer-based diffusion architecture, Vidu Q3 sets a new standard for cinematic and marketing vi
- Native Audio-Video Synchronization: Vidu Q3 generates lip-synced dialogue, sound effects, and background music simultaneously within a single pass, ensuring precise temporal alignment between audio tracks and visual lip movements without requiring post-processing.
- Extended High-Definition Video Generation: Supports up to 16 seconds of continuous video at 1080p resolution and 24 frames per second—the longest continuous generation duration among leading competitors—enabling more complex storytelling sequences.
- Smart Cuts for Scene Detection: Integrates automatic scene boundary detection and multi-shot narrative transitions, which facilitate the smooth generation of dynamic video scenes without manual intervention.
- Native Camera Control: Allows frame-level directorial commands such as pans, push-ins, and tracking shots within the generation pipeline, granting users granular cinematic control over the resulting video composition.
- Multimodal Input Flexibility: Accepts both text-to-video and image-to-video inputs with configurable start and end frame controls, enabling versatile use cases that range from scripted storyboarding to visual style transfer.
- Transformer-based Diffusion Architecture with Spatiotemporal Attention: The underlying Universal Vision Transformer (U-ViT) utilizes spatiotemporal attention mechanisms instead of conventional convolutional U-Nets, improving motion consistency and temporal coherence across generated frames.
Run Vidu Q3-Mix Reference to Video
from $0.16/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
images | Reference images for generating video with consistent subjects. Accepts 1 to 4 images as URLs or Base64 encoded strings. Supported codecs: PNG, JPEG, JPG, WebP. Dimensions must be at least 128x128 pixels, aspect ratio less than 1:4 or 4:1, | |
prompt | A textual description for video generation. | default Santa Claus and the bear hug by the lakeside. |
duration | The duration of the generated video in seconds. | default 5 |
resolution | The resolution of the generated media. Native 720p and 1080p use the original Vidu route when available. 1080p-sr generates a native 720p source video and applies FlashVSR super-resolution. 1440p-sr generates a native 1080p source video and | 720p 1080p 1080p-sr 1440p-sr |
generate_audio | Whether to generate audio for the video. | default True |
aspect_ratio | The aspect ratio of the output video. | 16:9 9:16 3:4 4:3 1:1 |
movement_amplitude | The movement amplitude of objects in the frame. | auto small medium large |
seed | The random seed to use for the generation. Set -1 for random. | default |
Sample prompt
The prompt behind the sample.
A massive futuristic spaceship flying through deep space, glowing engines leaving a subtle light trail, distant stars and nebula clouds in the background, cinematic lighting, dramatic scale, smooth camera tracking shot, ultra-realistic sci-fi, 4K
duration: 5resolution: 720pgenerate_audio: Trueaspect_ratio: 16:9movement_amplitude: autoseed: FAQ
Short answers.
How much does Vidu Q3-Mix Reference to Video cost?
Pricing starts at $0.16 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $1.27. Usage is billed per request from your balance — no subscription.
Does Vidu Q3-Mix Reference to Video run uncensored here?
No. Vidu applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Vidu Q3-Mix Reference to Video for everything else it does well.
What does Vidu Q3-Mix Reference to Video take as input?
It is a image-to-video model. Vidu Q3-Mix Reference-to-Video generates videos from 1-4 reference images with consistent subjects. Offers strong visual quality with intelligent scene transitions, smooth dynamic effects, and audio support up to 1080p.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related