Seedance 2.0 Mini Text-to-Video
Lightweight, economical video generation from text prompts with native audio.
What it does
Seedance 2.0 Mini Text-to-Video, in practice.
Seedance 2.0 Mini is the lightweight, cost-optimized tier of ByteDance's Seedance 2.0 family of multimodal generative AI models for synchronized video-and-audio creation. Developed by ByteDance and introduced on the CapCut/Dreamina platform in mid-2026, Mini inherits the same Dual-Branch Diffusion Transformer (DB-DiT) foundation and physics-informed world modeling as the flagship, but is tuned for throughput and price rather than maximum fidelity. Seedance 2.0 Mini renders at roughly twice the speed of Seedance 2.0 Fast while preserving comparable output quality, and it lowers generation cost by approximately 30% versus the standard Seedance 2.0 model (around half the cost at 720p). It keeps
- Most economical tier of the family: Mini is the lowest-cost Seedance 2.0 variant, designed for volume workloads where speed and price matter more than the last few percent of cinematic fidelity. It is roughly 2× faster than Seedance 2.0 Fast.
- Full multimodal input support: Like the rest of the family, Mini accepts text prompts, images, and reference videos/audio. The three exposed modes are text-to-video (prompt only), image-to-video (first frame, with optional last frame), and reference-to-video (multimodal references — images, video, and audio — for identity and style control).
- Native synchronized audio: Mini generates synchronized audio (voice, sound effects, and background music) alongside the video, built on the family's Dual-Branch Diffusion Transformer that couples the visual and audio streams.
- Reference-based consistency: The @ reference system preserves character identity and visual consistency across multiple generations, enabling coherent multi-shot sequences from a shared set of reference assets.
- World model with physics simulation: Mini retains the physics-informed world modeling of the Seedance 2.0 line, producing naturalistic object motion and stable spatial composition over the length of a clip.
- Tuned for iteration: Lower latency and lower cost make Mini well suited to storyboarding, A/B exploration of prompts, and producing many variants quickly before committing the best candidates to a higher-fidelity tier.
Run Seedance 2.0 Mini Text-to-Video
from $0.017/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | Text prompt describing the desired video. Supports Chinese and English. Recommended length: Chinese < 500 characters, English < 1000 words. | |
duration | Video duration in seconds (4-15), or -1 for model to choose automatically. | -1 4 5 6 7 8 9 10 11 12 13 14 |
resolution | Video resolution. | 480p 720p 720p-SR 1080p-SR 1440p-SR |
ratio | Aspect ratio. 'adaptive' automatically selects based on prompt context. | 16:9 4:3 1:1 3:4 9:16 21:9 adaptive |
bitrate_mode | Output video bitrate mode. 'high' encodes at a higher bitrate for a crisper, larger file; 'standard' uses the normal bitrate. Does not affect token cost. | standard high |
generate_audio | Whether to generate synchronized audio (voice, sound effects, background music). | default True |
seed | Seed integer used to control the randomness of generated content. Value range: [-1, 2^32-1]. The default -1 means a random seed is used. The same seed with the same request produces similar results, but complete consistency is not guarantee | default -1 |
watermark | Whether to add a watermark. | default |
return_last_frame | Whether to return the last frame as a separate image. | default |
Sample prompt
The prompt behind the sample.
A man in an elegant suit running across a wide, lush meadow while holding a bouquet of flowers, tall grass flowing in the wind, golden sunlight illuminating the landscape, joyful and emotional moment, cinematic composition, shallow depth of field, ultra-realistic, 4K, slow-motion tracking shot.
duration: 5resolution: 720pratio: adaptivegenerate_audio: Truewatermark: return_last_frame: FAQ
Short answers.
How much does Seedance 2.0 Mini Text-to-Video cost?
Pricing starts at $0.017 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.13. Usage is billed per request from your balance — no subscription.
Does Seedance 2.0 Mini Text-to-Video run uncensored here?
No. ByteDance applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Seedance 2.0 Mini Text-to-Video for everything else it does well.
What does Seedance 2.0 Mini Text-to-Video take as input?
It is a text-to-video model. Lightweight, economical video generation from text prompts with native audio.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related