Gemini Omni 1.1 Flash Video Edit
A natively multimodal Google DeepMind model that applies a text-instructed edit to an existing video - adding, removing, replacing, or restyling elements with native audio - while preserving everything the prompt does not mention.
What it does
Gemini Omni 1.1 Flash Video Edit, in practice.
Model ID: google/gemini-omni-1.1-flash/video-edit Gemini Omni 1.1 Flash is Google DeepMind's natively multimodal model for video generation and editing. This variant takes an existing video plus a text instruction and applies the edit — adding, removing, replacing, or restyling elements — while preserving everything the prompt does not mention.
- Flexible resolution output — new 4K and 1080p rendering for finished work, plus a fast 360p draft mode for previewing a shot before committing to a full-quality render.
- Longer continuous generation — shots can be extended segment by segment to a total of up to 40 seconds of coherent video.
- Precise shot start and end control — first and last frame can both be supplied, producing smoother camera moves, scene transitions, and seamless loops.
- Character and style consistency — a new video reference capability (clips of up to 3 seconds) substantially improves subject and art-direction stability.
- Instruction-driven editing — add, remove, replace, or restyle elements of a clip from a plain-language description.
- Scene-consistent results — edits blend into the existing footage, preserving untouched regions, lighting, and motion.
Run Gemini Omni 1.1 Flash Video Edit
from $0.055/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | Text prompt describing the edit to apply to the source video (e.g., add, remove, restyle, or transform elements). Maximum 20,000 characters. You can refer to an uploaded reference image by tag: <IMAGE_REF_N> is the Nth entry in reference_im | |
video | The source video to edit. Must be a publicly accessible URL, 10 seconds or less. Supported format: MP4. | |
reference_images | Images to use as character, scene, or style references for the edit. Accepts 1 to 10 images. Supported formats: PNG, JPEG, JPG, WebP. Each image is limited to 20MB. Supports both a public URL and a base64-encoded image for each item. | |
thinking_level | Controls the amount of internal reasoning the model performs before generating a response. Higher levels may improve quality on complex tasks but increase latency. | default high low |
seed | The random seed to use for the generation. -1 means a random seed will be used. | default -1 |
Sample prompt
The prompt behind the sample.
Convert the video style to an American animation style.
thinking_level: defaultseed: -1FAQ
Short answers.
How much does Gemini Omni 1.1 Flash Video Edit cost?
Pricing starts at $0.055 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.44. Usage is billed per request from your balance — no subscription.
Does Gemini Omni 1.1 Flash Video Edit run uncensored here?
No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Gemini Omni 1.1 Flash Video Edit for everything else it does well.
What does Gemini Omni 1.1 Flash Video Edit take as input?
It is a video-to-video model. A natively multimodal Google DeepMind model that applies a text-instructed edit to an existing video - adding, removing, replacing, or restyling elements with native audio - while preserving everything the prompt does not mention.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related