Gemini Omni Flash Text-to-Video
A natively multimodal Google DeepMind model that generates cinematic videos with synchronized native audio from a text prompt alone, grounded in real-world physics for controllable, high-speed video generation.
What it does
Gemini Omni Flash Text-to-Video, in practice.
Model ID: google/gemini-omni-flash/text-to-video Gemini Omni Flash is Google DeepMind's high-performance, natively multimodal model built for high-speed video generation, editing, and cinematic control. This variant accepts a text prompt only, making it ideal for pure creative generation where you describe the entire scene through language.
- Rich prompt understanding — Describe subjects, actions, camera movements, lighting, mood, style, and audio in a single prompt of up to 20,000 characters.
- Native audio generation — Every clip is rendered with a synchronized soundtrack (speech, music, effects) driven by your description.
- World-grounded realism — Physics, motion, and scene dynamics informed by Gemini's real-world knowledge.
- Cinematic control — Camera framing, pacing, and single-scene composition guided directly from the prompt.
- Adjustable reasoning — The `thinking_level` control trades latency for quality on complex prompts.
- Reproducible results — Set a fixed seed to reproduce or iterate on a specific generation.
Run Gemini Omni Flash Text-to-Video
from $0.19/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | Text prompt for generation. Describes the target content, style, camera language, or character actions. Maximum 20,000 characters. | |
duration | The duration of the generated video in seconds. | default 10 |
aspect_ratio | The aspect ratio of the generated video. | 16:9 9:16 |
resolution | The resolution of the generated video. | 720p |
thinking_level | Controls the amount of internal reasoning the model performs before generating a response. Higher levels may improve quality on complex tasks but increase latency. | default high low |
seed | The random seed to use for the generation. -1 means a random seed will be used. | default -1 |
Sample prompt
The prompt behind the sample.
A solitary man leans against a vintage red coupe on an empty seaside road at dusk. The wind gently moves the grass and his coat while a lonely cloud hangs motionless above him. The camera slowly tracks sideways, capturing reflections on the car body and the endless blue ocean beyond.
duration: 5aspect_ratio: 16:9resolution: 720pthinking_level: defaultseed: -1FAQ
Short answers.
How much does Gemini Omni Flash Text-to-Video cost?
Pricing starts at $0.19 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $1.50. Usage is billed per request from your balance — no subscription.
Does Gemini Omni Flash Text-to-Video run uncensored here?
No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Gemini Omni Flash Text-to-Video for everything else it does well.
What does Gemini Omni Flash Text-to-Video take as input?
It is a text-to-video model. A natively multimodal Google DeepMind model that generates cinematic videos with synchronized native audio from a text prompt alone, grounded in real-world physics for controllable, high-speed video generation.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related