Seedance 2.0 Fast Reference-to-Video
Fast multimodal video generation from reference images, videos, and audio. Supports video editing and extension.
What it does
Seedance 2.0 Fast Reference-to-Video, in practice.
Seedance 2.0 is a state-of-the-art multimodal generative AI model designed for synchronized video and audio content creation. Developed by ByteDance and integrated into the CapCut/Dreamina platform as of March 2026, this model family advances the field of generative multimedia by combining sophisticated diffusion transformer architectures with physics-informed world modeling for realistic motion and spatial consistency. Seedance 2.0’s significance lies in its Dual-Branch Diffusion Transformer (DB-DiT) architecture that jointly processes video and audio streams, enabling phoneme-level lip synchronization across multiple languages. Compared to previous iterations, it achieves substantially hig
- Dual-Branch Diffusion Transformer Architecture: Seedance 2.0 integrates separate yet synchronized diffusion branches for video and audio, enabling tight coupling between visual motion and sound generation. This architecture improves motion realism and audio-visual coherence beyond previous generative models.
- World Model with Physics Simulation: The model incorporates a physics-based world modeling approach that simulates realistic object motion and spatial consistency over time. This leads to naturalistic dynamics and stable scene composition across generated video sequences.
- Rich Multimodal Input Support: Seedance 2.0 accepts diverse input formats including text prompts, up to 9 images, and up to 3 video or audio clips of 15 seconds each. This flexibility allows nuanced content creation workflows combining static, dynamic, and auditory cues.
- Phoneme-Level Lip Synchronization: The native audio generation pipeline supports lip-sync at the phoneme granularity in 8+ languages, ensuring high fidelity mouth movements closely match generated speech or singing.
- High Usability and Efficiency: The model achieves an estimated 90% usable output rate compared to an industry average of approximately 20%, reducing post-processing overhead. Additionally, it delivers a 30% inference speed advantage over predecessor systems.
- API Variants for Different Use Cases: The Seedance 2.0 endpoint is geared toward high fidelity and cinematic visual effects suitable for final production, while the Seedance 2.0 Fast variant offers roughly 3 times faster generation and approximately 91% cost savings at $0.022 per second of output, ideal for rapid iteration and volume workflows.
Run Seedance 2.0 Fast Reference-to-Video
from $0.041/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | Text prompt describing the desired video. References like 'image 1', 'video 1' refer to inputs in order. | default The character in image 1 dances gracefully to the music |
reference_images | Reference image URLs, Base64, or asset references (asset://<ASSET_ID>). Up to 9 images for character/style/scene references, video editing, or combined generation. Per-image limits: formats jpeg/png/webp/bmp/tiff/gif/heic/heif, aspect ratio | |
reference_videos | Reference video URLs or asset references for video editing, extension, or multimodal generation. Up to 3 videos, total duration <= 15s. Per-video limits: formats mp4/mov, resolution 480p/720p/1080p, duration [2,15]s, aspect ratio (W/H) 0.4- | |
duration | Video duration in seconds (4-15), or -1 for model to choose automatically. | -1 4 5 6 7 8 9 10 11 12 13 14 |
resolution | Video resolution. | 480p 720p 720p-SR 1080p-SR 1440p-SR |
ratio | Aspect ratio. 'adaptive' uses primary media aspect ratio. | 16:9 4:3 1:1 3:4 9:16 21:9 adaptive |
bitrate_mode | Output video bitrate mode. 'high' encodes at a higher bitrate for a crisper, larger file; 'standard' uses the normal bitrate. Does not affect token cost. | standard high |
generate_audio | Whether to generate synchronized audio. | default True |
seed | Seed integer used to control the randomness of generated content. Value range: [-1, 2^32-1]. The default -1 means a random seed is used. The same seed with the same request produces similar results, but complete consistency is not guarantee | default -1 |
watermark | Whether to add a watermark. | default |
return_last_frame | Whether to return the last frame as a separate image. | default |
reference_audios | Reference audio URLs, Base64, or asset references. Must include at least 1 reference video or image. Formats: wav/mp3, duration [2,15]s, max 15MB each. Up to 3 audios, total duration <= 15s. |
Sample prompt
The prompt behind the sample.
Dark clouds rapidly transforming and racing across the sky above a majestic snow-covered mountain peak, dramatic high-altitude atmosphere, powerful winds shaping the clouds, cinematic wide shot, timelapse motion, volumetric lighting, epic nature documentary style, ultra-realistic, 4K.
duration: 5resolution: 720pratio: adaptivegenerate_audio: Truewatermark: return_last_frame: FAQ
Short answers.
How much does Seedance 2.0 Fast Reference-to-Video cost?
Pricing starts at $0.041 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.32. Usage is billed per request from your balance — no subscription.
Does Seedance 2.0 Fast Reference-to-Video run uncensored here?
No. ByteDance applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Seedance 2.0 Fast Reference-to-Video for everything else it does well.
What does Seedance 2.0 Fast Reference-to-Video take as input?
It is a image-to-video model. Fast multimodal video generation from reference images, videos, and audio. Supports video editing and extension.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related