Gemini Omni 1.1 Flash Video Extend
A natively multimodal Google DeepMind model that continues an existing clip with a seamlessly matched 3-to-10-second extension, chainable to grow a single shot up to a total of 40 seconds of coherent video with native audio.
What it does
Gemini Omni 1.1 Flash Video Extend, in practice.
Model ID: google/gemini-omni-1.1-flash/video-extend Gemini Omni 1.1 Flash is Google DeepMind's natively multimodal model for video generation and editing. This variant — new in 1.1 — takes an existing clip and continues it, appending a seamlessly matched 3-to-10-second continuation. Chain the calls and a single shot grows to a total of up to 40 seconds of coherent video.
- Flexible resolution output — new 4K and 1080p rendering for finished work, plus a fast 360p draft mode for previewing a shot before committing to a full-quality render.
- Longer continuous generation — scene extension analyses up to 10 seconds of prior context and grows a shot in increments, to a total of up to 40 seconds. This is the endpoint that exposes it.
- Precise shot start and end control — first and last frame can both be supplied, producing smoother camera moves, scene transitions, and seamless loops.
- Character and style consistency — a new video reference capability (clips of up to 3 seconds) substantially improves subject and art-direction stability.
- Seamless continuation — append 3 to 10 seconds that match the source's motion, lighting, subject, and audio across the join.
- Long-form assembly — chain extensions to build a single coherent shot of up to 40 seconds total.
Run Gemini Omni 1.1 Flash Video Extend
from $0.055/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | Text prompt describing how the source video should continue. Maximum 20,000 characters. You can refer to an uploaded reference image by tag: <IMAGE_REF_N> is the Nth entry in reference_images (0-based, so the first reference image is <IMAGE | |
video | The source video to extend. Must be a publicly accessible URL, 10 seconds or less. The model uses up to the last 10 seconds as context for the continuation. Supported format: MP4. | |
reference_images | Images to use as character, scene, or style references. Accepts 1 to 10 images. Supported formats: PNG, JPEG, JPG, WebP. Each image is limited to 20MB. Supports both a public URL and a base64-encoded image for each item. | |
duration | The duration in seconds of the continuation appended to the source video. The total length of the extended video must not exceed 40 seconds. | default 10 |
resolution | The resolution of the generated video. 360p is a fast, low-cost draft mode for previewing a shot; 1080p and 4k are upscaled from the natively generated frames. | 360p 720p 1080p 4k |
thinking_level | Controls the amount of internal reasoning the model performs before generating a response. Higher levels may improve quality on complex tasks but increase latency. | default high low |
seed | The random seed to use for the generation. -1 means a random seed will be used. | default -1 |
Sample prompt
The prompt behind the sample.
Zooming in further on the video reveals the forest shown in the image.
duration: 10resolution: 720pthinking_level: defaultseed: -1FAQ
Short answers.
How much does Gemini Omni 1.1 Flash Video Extend cost?
Pricing starts at $0.055 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.44. Usage is billed per request from your balance — no subscription.
Does Gemini Omni 1.1 Flash Video Extend run uncensored here?
No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Gemini Omni 1.1 Flash Video Extend for everything else it does well.
What does Gemini Omni 1.1 Flash Video Extend take as input?
It is a video-to-video model. A natively multimodal Google DeepMind model that continues an existing clip with a seamlessly matched 3-to-10-second extension, chainable to grow a single shot up to a total of 40 seconds of coherent video with native audio.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related