Gemini Omni 1.1 Flash Image-to-Video
A natively multimodal Google DeepMind model that animates a still image into a cinematic, natively sound-enabled clip from a text prompt, optionally interpolating to a supplied last frame for precise shot start and end control.
What it does
Gemini Omni 1.1 Flash Image-to-Video, in practice.
Model ID: google/gemini-omni-1.1-flash/image-to-video Gemini Omni 1.1 Flash is Google DeepMind's natively multimodal model for video generation and editing. This variant animates a still image according to a text prompt, and — new in 1.1 — accepts an optional last frame so the model interpolates the whole shot between two keyframes you choose.
- Flexible resolution output — new 4K and 1080p rendering for finished work, plus a fast 360p draft mode for previewing a shot before committing to a full-quality render.
- Longer continuous generation — shots can be extended segment by segment to a total of up to 40 seconds of coherent video.
- Precise shot start and end control — first and last frame can both be supplied, producing smoother camera moves, scene transitions, and seamless loops.
- Character and style consistency — a new video reference capability (clips of up to 3 seconds) substantially improves subject and art-direction stability.
- Image animation — bring a photograph, render, or illustration into motion while preserving its subject and composition.
- First / last frame interpolation — supply both ends of the shot and let the model generate a smooth path between them.
Run Gemini Omni 1.1 Flash Image-to-Video
from $0.058/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | Text prompt for generation. Describes the target content, style, camera language, or character actions. Maximum 20,000 characters. | |
image | The first frame of the generated video. Supported formats: PNG, JPEG, JPG, WebP. Limited to 20MB. Supports both a public URL and a base64-encoded image. | |
last_image | The last frame of the generated video. The model interpolates the shot between image and last_image. Requires image to be set. Supported formats: PNG, JPEG, JPG, WebP. Limited to 20MB. Supports both a public URL and a base64-encoded image. | |
duration | The duration of the generated video in seconds. | default 10 |
aspect_ratio | The aspect ratio of the generated video. | 16:9 9:16 |
resolution | The resolution of the generated video. 360p is a fast, low-cost draft mode for previewing a shot; 1080p and 4k are upscaled from the natively generated frames. | 360p 720p 1080p 4k |
thinking_level | Controls the amount of internal reasoning the model performs before generating a response. Higher levels may improve quality on complex tasks but increase latency. | default high low |
seed | The random seed to use for the generation. -1 means a random seed will be used. | default -1 |
Sample prompt
The prompt behind the sample.
A colossal ancient stone arch stands alone in an immense alien desert, towering above the landscape like a forgotten monument from an ancient civilization. The camera begins with a wide establishing shot, slowly pushing forward toward the enormous arch while a lone human figure stands motionless in the foreground, emphasizing the overwhelming sense of scale. Massive clouds drift slowly across the blue-gray sky, their shadows moving gently across the desert. Fine dust and sand particles are carried by the wind around the base of the monument. As the camera approaches, subtle light breaks throug
duration: 8aspect_ratio: 9:16resolution: 720pthinking_level: defaultseed: -1FAQ
Short answers.
How much does Gemini Omni 1.1 Flash Image-to-Video cost?
Pricing starts at $0.058 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.47. Usage is billed per request from your balance — no subscription.
Does Gemini Omni 1.1 Flash Image-to-Video run uncensored here?
No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Gemini Omni 1.1 Flash Image-to-Video for everything else it does well.
What does Gemini Omni 1.1 Flash Image-to-Video take as input?
It is a image-to-video model. A natively multimodal Google DeepMind model that animates a still image into a cinematic, natively sound-enabled clip from a text prompt, optionally interpolating to a supplied last frame for precise shot start and end control.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related