Gemini Omni 1.1 Flash Text-to-Video
A natively multimodal Google DeepMind model that turns a single text prompt into a cinematic clip with synchronized native audio, with control over duration, aspect ratio, and output resolution from a fast 360p draft up to 4K.
What it does
Gemini Omni 1.1 Flash Text-to-Video, in practice.
Model ID: google/gemini-omni-1.1-flash/text-to-video Gemini Omni 1.1 Flash is Google DeepMind's natively multimodal model for video generation and editing. This variant turns a single text prompt into a fully rendered clip — picture and synchronized audio together — with control over duration, aspect ratio, and output resolution from a 360p draft all the way up to 4K.
- Flexible resolution output — new 4K and 1080p rendering for finished work, plus a fast 360p draft mode for previewing a shot before committing to a full-quality render.
- Longer continuous generation — shots can be extended segment by segment to a total of up to 40 seconds of coherent video.
- Precise shot start and end control — first and last frame can both be supplied, producing smoother camera moves, scene transitions, and seamless loops.
- Character and style consistency — a new video reference capability (clips of up to 3 seconds) substantially improves subject and art-direction stability.
- Text-only generation — a complete audiovisual clip from a single written description, no reference media required.
- Native audio generation — every clip is rendered with a synchronized soundtrack (speech, music, effects) driven by your description.
Run Gemini Omni 1.1 Flash Text-to-Video
from $0.055/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | Text prompt for generation. Describes the target content, style, camera language, or character actions. Maximum 20,000 characters. | |
duration | The duration of the generated video in seconds. | default 10 |
aspect_ratio | The aspect ratio of the generated video. | 16:9 9:16 |
resolution | The resolution of the generated video. 360p is a fast, low-cost draft mode for previewing a shot; 1080p and 4k are upscaled from the natively generated frames. | 360p 720p 1080p 4k |
thinking_level | Controls the amount of internal reasoning the model performs before generating a response. Higher levels may improve quality on complex tasks but increase latency. | default high low |
seed | The random seed to use for the generation. -1 means a random seed will be used. | default -1 |
Sample prompt
The prompt behind the sample.
A cinematic sci-fi spacecraft races through a massive golden cosmic vortex, following the curvature of an enormous glowing ring. The spacecraft accelerates rapidly into the distance, leaving long trails of white-blue engine light behind it. Golden energy streams and luminous particles sweep past the camera at extreme speed, creating powerful parallax and a strong sense of motion. The camera dynamically tracks alongside and slightly behind the spacecraft, then slowly pushes forward as the ship dives deeper into the vortex. Massive arcs of golden light bend across the dark universe, while clouds
duration: 7aspect_ratio: 16:9resolution: 720pthinking_level: defaultseed: -1FAQ
Short answers.
How much does Gemini Omni 1.1 Flash Text-to-Video cost?
Pricing starts at $0.055 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.44. Usage is billed per request from your balance — no subscription.
Does Gemini Omni 1.1 Flash Text-to-Video run uncensored here?
No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Gemini Omni 1.1 Flash Text-to-Video for everything else it does well.
What does Gemini Omni 1.1 Flash Text-to-Video take as input?
It is a text-to-video model. A natively multimodal Google DeepMind model that turns a single text prompt into a cinematic clip with synchronized native audio, with control over duration, aspect ratio, and output resolution from a fast 360p draft up to 4K.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related