Seedance 2.5 Reference-to-Video
Multimodal video generation from reference images, videos, and audio. Supports video editing and extension.
What it does
Seedance 2.5 Reference-to-Video, in practice.
Seedance 2.5 is ByteDance's next-generation multimodal video generation model, officially unveiled by Volcano Engine president Tan Dai at the 2026 Volcano Engine FORCE conference in Beijing on June 23, 2026. This README covers the following API model identifiers: Succeeding the Seedance 2.0 family, Seedance 2.5 advances generative video along four axes announced at launch: native single-pass generation of clips up to 30 seconds (double the 15-second ceiling of Seedance 2.0, with substantially improved shot-to-shot camera continuity), joint conditioning on up to 50 all-modality reference assets (up from 12), precise consistency-preserving video editing and extension, and native multilingual g
- 30-Second Single-Pass Generation: Produces a complete video of up to 30 seconds in one native generation pass — no stitching of shorter segments — with markedly improved camera and shot continuity across the full clip. This doubles the Seedance 2.0 family's 15-second output ceiling and enables genuine short-narrative work in a single request.
- 50 All-Modality Reference Assets: The reference-to-video variant conditions jointly on up to 50 reference materials — up to 30 reference images, 10 reference videos, and 10 reference audios in a single request, with a combined audio/video reference budget of 30 seconds. That expands Seedance 2.0's capacity (9 images and 3 audio/video clips, 15 seconds total) on every axis. Subjects, styles, motion cues, and audio timing can all be anchored to user-supplied assets at once.
- Audio-Only References: New in this release, a single BGM track, voice track, or sound-effect track can serve as the sole reference — directly guiding visual pacing, beat matching, and lip synchronization without any accompanying image or video input.
- Consistency-Preserving Localized Editing and Extension: Introduces precise video editing that keeps the overall frame intact while changing only targeted local elements, alongside high-fidelity temporal extension of existing footage — extending the family's world-model approach from pure generation into controllable editing workflows.
- Three Task-Focused Variants: `text-to-video` generates from a prompt alone; `image-to-video` animates a first frame (optionally pinning a last frame for precise start/end control); `reference-to-video` composes new footage from large multimodal reference sets. All variants share the same generation core, prompt understanding, and audio pipeline.
- Flexible Duration and Output Controls: Requests specify any duration from 4 to 30 seconds (or delegate the choice to the model), native 480p, 720p, or 1080p generation, optional FlashVSR `-sr` and video enhance `-esr` delivery modes up to 4K, aspect ratios from vertical 9:16 to widescreen 16:9, MP4 or MOV containers, plus seed, watermark, and audio-generation toggles for reproducible, pipeline-ready results.
Run Seedance 2.5 Reference-to-Video
from $0.20/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | Text prompt describing the desired video. Cite reference inputs in submission order with @-syntax: @Image1, @Video1, @Audio1, etc. Prompts phrased as EDITING an input video (e.g. modifying/continuing @Video1's own scene) switch the model in | default The character in image 1 dances gracefully to the music |
reference_images | Reference image URLs, Base64, or asset references (asset://<ASSET_ID>). Up to 30 images for character/style/scene references, video editing, or combined generation. Per-image limits: formats jpeg/png/webp/bmp/tiff/gif/heic/heif, aspect rati | |
reference_videos | Reference video URLs or asset references for multimodal reference, video editing, or extension. Up to 10 videos. Per-video limits: formats mp4/mov (H.264/H.265 + AAC/MP3), resolution 480p-4k, duration [2,30]s, aspect ratio (W/H) 0.4-2.5, wi | |
duration | Video duration in seconds (4-30), or -1 for the model to choose automatically. Video-EDITING prompts (operating on an input video) accept only -1: the output tracks the input video's length (may be ~0.4s shorter). | -1 4 5 6 7 8 9 10 11 12 13 14 |
resolution | Video resolution. 480p, 720p, and 1080p are native Seedance outputs. Every -sr and -esr option first generates the nearest native source, then upscales or enhances it: 720p-sr and 720p-esr from 480p; 1080p-sr, 1080p-esr, and 1080p-esr & 60f | 480p 720p 720p-sr 720p-esr 1080p 1080p-sr 1080p-esr 1080p-esr & 60fps 1440p-sr 1440p-esr 4k-esr |
ratio | Aspect ratio. 'adaptive' uses the primary media aspect ratio. Explicit ratios apply to reference-to-video generation; video editing and extension force 'adaptive' (the source video's ratio is preserved). | 16:9 4:3 1:1 3:4 9:16 21:9 adaptive |
generate_audio | Whether to generate synchronized audio. | default True |
watermark | Whether to add a watermark. | default |
return_last_frame | Whether to return the last frame as a separate image. | default |
reference_audios | Reference audio URLs, Base64, or asset references. Formats: wav/mp3, duration [2,30]s per clip, max 15MB each. Up to 10 audios; combined duration of all reference audios must not exceed 30s per request. Audio-only referencing is supported ( | |
output_format | Output video container format. 'mp4' is the default; 'mov' encodes yuv444p for higher color fidelity, suited to multi-round editing/extension pipelines where recompression loss accumulates. | mp4 mov |
Sample prompt
The prompt behind the sample.
On the tenth year of the Trojan War, beneath a moonlit sky heavy with salt and ash, Odysseus@image2 unveils his wooden horse stratagem to four kings beside its half-built skeleton. Achilles@image1 , blood still drying on his bronze armor, scorns deception as a shepherd's trick unworthy of his spear; Ajax the Great@image3 stands with him, his tower shield casting a shadow over the sand, demanding honorable battle. Menelaus@image5 , eyes burning with a decade of humiliation, grips his sword and declares he cares nothing for method—only that the gates of Troy open. Between pride and vengeance,
duration: 5resolution: 720pratio: 16:9generate_audio: Truewatermark: return_last_frame: output_format: mp4FAQ
Short answers.
How much does Seedance 2.5 Reference-to-Video cost?
Pricing starts at $0.20 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $1.61. Usage is billed per request from your balance — no subscription.
Does Seedance 2.5 Reference-to-Video run uncensored here?
No. ByteDance applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Seedance 2.5 Reference-to-Video for everything else it does well.
What does Seedance 2.5 Reference-to-Video take as input?
It is a image-to-video model. Multimodal video generation from reference images, videos, and audio. Supports video editing and extension.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related