Gemini Omni Flash Video Edit
A natively multimodal Google DeepMind model that edits an existing video from a text prompt with optional reference images, applying scene-consistent changes and native audio while preserving the untouched footage.
What it does
Gemini Omni Flash Video Edit, in practice.
Model ID: google/gemini-omni-flash/video-edit Gemini Omni Flash is Google DeepMind's high-performance, natively multimodal model built for high-speed video generation, editing, and cinematic control. This variant accepts a source video plus a text prompt (and, optionally, reference images), transforming an existing clip according to your instructions.
- Instruction-driven editing — Add, remove, replace, or restyle elements of a clip from a plain-language description.
- Scene-consistent results — Edits blend into the existing footage, preserving untouched regions, lighting, and motion.
- Reference-guided edits — Optionally supply up to 5 images to introduce a specific subject, object, or style.
- Native audio generation — Edits can regenerate or adjust the accompanying soundtrack alongside the picture.
- World-grounded realism — Physics, motion, and scene dynamics informed by Gemini's real-world knowledge.
- Adjustable reasoning — The `thinking_level` control trades latency for quality on complex edits.
Run Gemini Omni Flash Video Edit
from $0.21/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
video | The source video to edit. Limited to 100MB and 30 seconds duration. | |
prompt | Text prompt describing the edit to apply to the source video (e.g., add, remove, or transform elements). Maximum 20,000 characters. | |
images | Images to use as character, scene, or style references. Accepts 1 to 10 images when combined with a video reference. Supported formats: PNG, JPEG, JPG, WebP. Each image is limited to 20MB. Supports both a public URL and a base64-encoded ima | |
resolution | The resolution of the generated video. | 720p |
thinking_level | Controls the amount of internal reasoning the model performs before generating a response. Higher levels may improve quality on complex tasks but increase latency. | default high low |
seed | The random seed to use for the generation. -1 means a random seed will be used. | default -1 |
Sample prompt
The prompt behind the sample.
Change the overall color of the headphones to red.
resolution: 720pthinking_level: defaultseed: -1FAQ
Short answers.
How much does Gemini Omni Flash Video Edit cost?
Pricing starts at $0.21 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $1.68. Usage is billed per request from your balance — no subscription.
Does Gemini Omni Flash Video Edit run uncensored here?
No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Gemini Omni Flash Video Edit for everything else it does well.
What does Gemini Omni Flash Video Edit take as input?
It is a video-to-video model. A natively multimodal Google DeepMind model that edits an existing video from a text prompt with optional reference images, applying scene-consistent changes and native audio while preserving the untouched footage.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related