Gemini Omni Flash Image-to-Video
A natively multimodal Google DeepMind model that animates a still image into a cinematic, sound-enabled video guided by a text prompt while preserving the source subject and composition.
What it does
Gemini Omni Flash Image-to-Video, in practice.
Model ID: google/gemini-omni-flash/image-to-video Gemini Omni Flash is Google DeepMind's high-performance, natively multimodal model built for high-speed video generation, editing, and cinematic control. This variant accepts an image plus a text prompt, animating a still image into a coherent, sound-enabled video guided by your instructions.
- Image-grounded animation — Preserves the subject, style, and composition of the source image while adding motion.
- Rich prompt understanding — Direct camera movement, action, mood, style, and audio in a single prompt of up to 20,000 characters.
- Native audio generation — Every clip is rendered with a synchronized soundtrack (speech, music, effects) driven by your description.
- World-grounded realism — Physics, motion, and scene dynamics informed by Gemini's real-world knowledge.
- Adjustable reasoning — The `thinking_level` control trades latency for quality on complex prompts.
- Reproducible results — Set a fixed seed to reproduce or iterate on a specific generation.
Run Gemini Omni Flash Image-to-Video
from $0.20/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
image | The image to animate into a video, used as the starting frame or motion guide. Supported formats: PNG, JPEG, JPG, WebP. Limited to 20MB. Supports both a public URL and a base64-encoded image. | |
prompt | Text prompt for generation. Describes the target content, style, camera language, or character actions. Maximum 20,000 characters. | |
duration | The duration of the generated video in seconds. | default 10 |
aspect_ratio | The aspect ratio of the generated video. | 16:9 9:16 |
resolution | The resolution of the generated video. | 720p |
thinking_level | Controls the amount of internal reasoning the model performs before generating a response. Higher levels may improve quality on complex tasks but increase latency. | default high low |
seed | The random seed to use for the generation. -1 means a random seed will be used. | default -1 |
Sample prompt
The prompt behind the sample.
The spacecraft is flying at high speed.
duration: 5aspect_ratio: 16:9resolution: 720pthinking_level: defaultseed: -1FAQ
Short answers.
How much does Gemini Omni Flash Image-to-Video cost?
Pricing starts at $0.20 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $1.56. Usage is billed per request from your balance — no subscription.
Does Gemini Omni Flash Image-to-Video run uncensored here?
No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Gemini Omni Flash Image-to-Video for everything else it does well.
What does Gemini Omni Flash Image-to-Video take as input?
It is a image-to-video model. A natively multimodal Google DeepMind model that animates a still image into a cinematic, sound-enabled video guided by a text prompt while preserving the source subject and composition.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related