minimaxiH3 Get access
ViduImage-to-videoVendor content policy applies

Vidu Q3 Reference to Video

Vidu Q3 Reference-to-Video generates videos from 1-4 reference images with consistent subjects. Features intelligent camera switching with better consistency across multiple camera positions, audio support, and resolutions up to 1080p.

$0.063per secondStarting price at the base resolution and quality tier.
$0.50Typical 8-second clip at the base tier. Longer or higher-resolution outputs scale with duration and tier.
Pay per useBilled per request from your balance. No subscription, no minimum.
Sample output · prompt: “A lone figure walking slowly toward a massive futuristic structure in the distance, vast sci-fi landscape, glowing architectural lights, cinematic atmosphere, soft fog, dramatic scale, smooth camera tracking shot, ultra-…”

What it does

Vidu Q3 Reference to Video, in practice.

Vidu Q3 is an advanced AI video generation model developed by Shengshu Technology (生数科技) in collaboration with Tsinghua University. Released on January 30, 2026, Vidu Q3 is designed to produce high-fidelity, synchronized audio-visual content with industry-leading continuous video length and native support for integrated audio generation. The model represents a significant advancement in automated video synthesis by unifying multiple complex video generation tasks—such as lip-synced dialogue, dynamic camera movements, and multi-shot storytelling—into a single-pass framework. Leveraging a novel Transformer-based diffusion architecture, Vidu Q3 sets a new standard for cinematic and marketing vi

  • Native Audio-Video Synchronization: Vidu Q3 generates lip-synced dialogue, sound effects, and background music simultaneously within a single pass, ensuring precise temporal alignment between audio tracks and visual lip movements without requiring post-processing.
  • Extended High-Definition Video Generation: Supports up to 16 seconds of continuous video at 1080p resolution and 24 frames per second—the longest continuous generation duration among leading competitors—enabling more complex storytelling sequences.
  • Smart Cuts for Scene Detection: Integrates automatic scene boundary detection and multi-shot narrative transitions, which facilitate the smooth generation of dynamic video scenes without manual intervention.
  • Native Camera Control: Allows frame-level directorial commands such as pans, push-ins, and tracking shots within the generation pipeline, granting users granular cinematic control over the resulting video composition.
  • Multimodal Input Flexibility: Accepts both text-to-video and image-to-video inputs with configurable start and end frame controls, enabling versatile use cases that range from scripted storyboarding to visual style transfer.
  • Transformer-based Diffusion Architecture with Spatiotemporal Attention: The underlying Universal Vision Transformer (U-ViT) utilizes spatiotemporal attention mechanisms instead of conventional convolutional U-Nets, improving motion consistency and temporal coherence across generated frames.

Run Vidu Q3 Reference to Video

from $0.063/sec
Drop an image hereor click to choose a fileFiles stay in your project. Nothing is trained on.
resolution
aspect_ratio
movement_amplitude
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
imagesReference images for generating video with consistent subjects. Accepts 1 to 4 images as URLs or Base64 encoded strings. Supported codecs: PNG, JPEG, JPG, WebP. Dimensions must be at least 128x128 pixels, aspect ratio less than 1:4 or 4:1,
promptA textual description for video generation.default Santa Claus and the bear hug by the lakeside.
durationThe duration of the generated video in seconds.default 5
resolutionThe resolution of the generated media. Native 540p, 720p, and 1080p use the original Vidu route when available. 1080p-sr generates a native 720p source video and applies FlashVSR super-resolution. 1440p-sr generates a native 1080p source vi540p 720p 1080p 1080p-sr 1440p-sr
generate_audioWhether to generate audio for the video.default True
aspect_ratioThe aspect ratio of the output video.16:9 9:16 3:4 4:3 1:1
movement_amplitudeThe movement amplitude of objects in the frame.auto small medium large
seedThe random seed to use for the generation. Set -1 for random.default

Sample prompt

The prompt behind the sample.

A lone figure walking slowly toward a massive futuristic structure in the distance, vast sci-fi landscape, glowing architectural lights, cinematic atmosphere, soft fog, dramatic scale, smooth camera tracking shot, ultra-detailed, 4K
duration: 5resolution: 720pgenerate_audio: Trueaspect_ratio: 16:9movement_amplitude: autoseed:

FAQ

Short answers.

How much does Vidu Q3 Reference to Video cost?

Pricing starts at $0.063 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.50. Usage is billed per request from your balance — no subscription.

Does Vidu Q3 Reference to Video run uncensored here?

No. Vidu applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Vidu Q3 Reference to Video for everything else it does well.

What does Vidu Q3 Reference to Video take as input?

It is a image-to-video model. Vidu Q3 Reference-to-Video generates videos from 1-4 reference images with consistent subjects. Features intelligent camera switching with better consistency across multiple camera positions, audio support, and resolutions up to 1080p.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run Vidu Q3 Reference to Video.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models