Gemini Omni Flash Reference-to-Video
A natively multimodal Google DeepMind model that generates cinematic, sound-enabled videos from a text prompt plus 1-5 reference images, carrying a consistent subject, scene, or style across generations.
What it does
Gemini Omni Flash Reference-to-Video, in practice.
Model ID: google/gemini-omni-flash/reference-to-video Gemini Omni Flash is Google DeepMind's high-performance, natively multimodal model built for high-speed video generation, editing, and cinematic control. This variant accepts a text prompt plus one or more reference images, generating a video that carries the referenced subject, scene, or style into a newly described scene.
- Subject & style consistency — Carry a referenced character, object, or look across scenes and generations.
- Multi-reference conditioning — Blend up to 5 reference images to guide subject, scene, and style at once.
- Rich prompt understanding — Direct camera movement, action, mood, style, and audio in a single prompt of up to 20,000 characters.
- Native audio generation — Every clip is rendered with a synchronized soundtrack (speech, music, effects) driven by your description.
- World-grounded realism — Physics, motion, and scene dynamics informed by Gemini's real-world knowledge.
- Adjustable reasoning — The `thinking_level` control trades latency for quality on complex prompts.
Run Gemini Omni Flash Reference-to-Video
from $0.20/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | Text prompt for generation. Describes the target content, style, camera language, or character actions. Maximum 20,000 characters. | |
images | Images to use as character, scene, or style references. Accepts 1 to 10 images when combined with a video reference. Supported formats: PNG, JPEG, JPG, WebP. Each image is limited to 20MB. Supports both a public URL and a base64-encoded ima | |
duration | The duration of the generated video in seconds. | default 10 |
aspect_ratio | The aspect ratio of the generated video. | 16:9 9:16 |
resolution | The resolution of the generated video. | 720p |
thinking_level | Controls the amount of internal reasoning the model performs before generating a response. Higher levels may improve quality on complex tasks but increase latency. | default high low |
seed | The random seed to use for the generation. -1 means a random seed will be used. | default -1 |
Sample prompt
The prompt behind the sample.
The little doll is jumping happily.
duration: 10aspect_ratio: 16:9resolution: 720pthinking_level: defaultseed: -1FAQ
Short answers.
How much does Gemini Omni Flash Reference-to-Video cost?
Pricing starts at $0.20 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $1.62. Usage is billed per request from your balance — no subscription.
Does Gemini Omni Flash Reference-to-Video run uncensored here?
No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Gemini Omni Flash Reference-to-Video for everything else it does well.
What does Gemini Omni Flash Reference-to-Video take as input?
It is a text-to-video model. A natively multimodal Google DeepMind model that generates cinematic, sound-enabled videos from a text prompt plus 1-5 reference images, carrying a consistent subject, scene, or style across generations.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related