minimaxiH3 Get access
ByteDanceImage-to-videoVendor content policy applies

Seedance 2.0 Mini Reference-to-Video

Lightweight, economical multimodal video generation from reference images, videos, and audio with native audio.

$0.017per secondStarting price at the base resolution and quality tier.
$0.13Typical 8-second clip at the base tier. Longer or higher-resolution outputs scale with duration and tier.
Pay per useBilled per request from your balance. No subscription, no minimum.
Sample output · prompt: “Car 1 @image1 is speeding along the highway@image3 , while Car 2@image2 , with its hazard lights flashing, is rapidly closing in from the rear left.”

What it does

Seedance 2.0 Mini Reference-to-Video, in practice.

Seedance 2.0 Mini is the lightweight, cost-optimized tier of ByteDance's Seedance 2.0 family of multimodal generative AI models for synchronized video-and-audio creation. Developed by ByteDance and introduced on the CapCut/Dreamina platform in mid-2026, Mini inherits the same Dual-Branch Diffusion Transformer (DB-DiT) foundation and physics-informed world modeling as the flagship, but is tuned for throughput and price rather than maximum fidelity. Seedance 2.0 Mini renders at roughly twice the speed of Seedance 2.0 Fast while preserving comparable output quality, and it lowers generation cost by approximately 30% versus the standard Seedance 2.0 model (around half the cost at 720p). It keeps

  • Most economical tier of the family: Mini is the lowest-cost Seedance 2.0 variant, designed for volume workloads where speed and price matter more than the last few percent of cinematic fidelity. It is roughly 2× faster than Seedance 2.0 Fast.
  • Full multimodal input support: Like the rest of the family, Mini accepts text prompts, images, and reference videos/audio. The three exposed modes are text-to-video (prompt only), image-to-video (first frame, with optional last frame), and reference-to-video (multimodal references — images, video, and audio — for identity and style control).
  • Native synchronized audio: Mini generates synchronized audio (voice, sound effects, and background music) alongside the video, built on the family's Dual-Branch Diffusion Transformer that couples the visual and audio streams.
  • Reference-based consistency: The @ reference system preserves character identity and visual consistency across multiple generations, enabling coherent multi-shot sequences from a shared set of reference assets.
  • World model with physics simulation: Mini retains the physics-informed world modeling of the Seedance 2.0 line, producing naturalistic object motion and stable spatial composition over the length of a clip.
  • Tuned for iteration: Lower latency and lower cost make Mini well suited to storyboarding, A/B exploration of prompts, and producing many variants quickly before committing the best candidates to a higher-fidelity tier.

Run Seedance 2.0 Mini Reference-to-Video

from $0.017/sec
Drop an image hereor click to choose a fileFiles stay in your project. Nothing is trained on.
resolution
ratio
bitrate_mode
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
promptText prompt describing the desired video. References like 'image 1', 'video 1' refer to inputs in order.default The character in image 1 dances gracefully to the music
reference_imagesReference image URLs, Base64, or asset references (asset://<ASSET_ID>). Up to 9 images for character/style/scene references, video editing, or combined generation. Per-image limits: formats jpeg/png/webp/bmp/tiff/gif/heic/heif, aspect ratio
reference_videosReference video URLs or asset references for video editing, extension, or multimodal generation. Up to 3 videos, total duration <= 15s. Per-video limits: formats mp4/mov, resolution 480p/720p/1080p, duration [2,15]s, aspect ratio (W/H) 0.4-
durationVideo duration in seconds (4-15), or -1 for model to choose automatically.-1 4 5 6 7 8 9 10 11 12 13 14
resolutionVideo resolution.480p 720p 720p-SR 1080p-SR 1440p-SR
ratioAspect ratio. 'adaptive' uses primary media aspect ratio.16:9 4:3 1:1 3:4 9:16 21:9 adaptive
bitrate_modeOutput video bitrate mode. 'high' encodes at a higher bitrate for a crisper, larger file; 'standard' uses the normal bitrate. Does not affect token cost.standard high
generate_audioWhether to generate synchronized audio.default True
seedSeed integer used to control the randomness of generated content. Value range: [-1, 2^32-1]. The default -1 means a random seed is used. The same seed with the same request produces similar results, but complete consistency is not guaranteedefault -1
watermarkWhether to add a watermark.default
return_last_frameWhether to return the last frame as a separate image.default
reference_audiosReference audio URLs, Base64, or asset references. Must include at least 1 reference video or image. Formats: wav/mp3, duration [2,15]s, max 15MB each. Up to 3 audios, total duration <= 15s.

Sample prompt

The prompt behind the sample.

Car 1 @image1 is speeding along the highway@image3 , while Car 2@image2 , with its hazard lights flashing, is rapidly closing in from the rear left.
duration: 5resolution: 720pratio: 16:9generate_audio: Truewatermark: return_last_frame: bitrate_mode: standardseed: -1

FAQ

Short answers.

How much does Seedance 2.0 Mini Reference-to-Video cost?

Pricing starts at $0.017 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.13. Usage is billed per request from your balance — no subscription.

Does Seedance 2.0 Mini Reference-to-Video run uncensored here?

No. ByteDance applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Seedance 2.0 Mini Reference-to-Video for everything else it does well.

What does Seedance 2.0 Mini Reference-to-Video take as input?

It is a image-to-video model. Lightweight, economical multimodal video generation from reference images, videos, and audio with native audio.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run Seedance 2.0 Mini Reference-to-Video.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models