Gemini Omni Flash Image-to-Video Developer
Gemini Omni Flash is Google's multimodal video generation model. This image-to-video variant creates subject-consistent videos from up to 7 reference images combined with a text prompt, preserving visual identity across the full generated video.
What it does
Gemini Omni Flash Image-to-Video Developer, in practice.
Model ID: google/gemini-omni-flash/image-to-video-developer Gemini Omni is Google's multimodal video generation model designed to create high-quality video content from diverse input types. This variant accepts a text prompt plus up to 7 reference images, enabling subject-consistent video generation where the visual identity of characters, objects, or scenes is anchored by real image references.
- Image-guided generation — Provide 1 to 7 reference images to anchor subjects, environments, or visual styles.
- Subject consistency — The model preserves key visual details from the reference images across the full video duration.
- Rich prompt understanding — Complement image references with a prompt of up to 20,000 characters describing actions, camera movements, lighting, and mood.
- Multi-resolution output — Generate at 720p, 1080p, or 4K.
- Flexible aspect ratios — 16:9 landscape or 9:16 portrait.
- Controllable duration — 4, 6, 8, or 10 seconds per generation.
Run Gemini Omni Flash Image-to-Video Developer
from $0.17/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | Text prompt for generation. Describes the target content, style, camera language, or character actions. Maximum 20,000 characters. | |
images | Images to use as character, scene, or style references. Accepts 1 to 7 images. Supported formats: PNG, JPEG, JPG, WebP. Each image is limited to 20MB. | |
duration | The duration of the generated video in seconds. | 4 6 8 10 |
aspect_ratio | The aspect ratio of the generated video. | 16:9 9:16 |
resolution | The resolution of the generated video. | 720p 1080p 4k |
seed | Random seed for reproducibility. Use -1 to use a random seed. | default -1 |
Sample prompt
The prompt behind the sample.
Style & Camera: High-speed cinematic action photography, 4K resolution, shallow depth of field. The camera begins with a low-angle tracking shot, moving parallel to a soccer player sprinting down a rain-slicked green pitch under dramatic stadium floodlights. Visual & Action Sequence: The soccer player, wearing a detailed jersey with realistic fabric ripples, powerful plants their left foot into the wet turf. Water droplets and blades of grass explode upwards from the impact. As they swing their right leg back and powerfully strike the soccer ball, the ball deforms slightly against the boot be
duration: 6aspect_ratio: 16:9resolution: 720pseed: -1FAQ
Short answers.
How much does Gemini Omni Flash Image-to-Video Developer cost?
Pricing starts at $0.17 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $1.34. Usage is billed per request from your balance — no subscription.
Does Gemini Omni Flash Image-to-Video Developer run uncensored here?
No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Gemini Omni Flash Image-to-Video Developer for everything else it does well.
What does Gemini Omni Flash Image-to-Video Developer take as input?
It is a image-to-video model. Gemini Omni Flash is Google's multimodal video generation model. This image-to-video variant creates subject-consistent videos from up to 7 reference images combined with a text prompt, preserving visual identity across the full generated video.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related