Gemini Omni 1.1 Flash Reference-to-Video
A natively multimodal Google DeepMind model that generates cinematic, natively sound-enabled videos from a text prompt plus up to 10 reference images and 3 reference video clips, keeping a character, product, or art direction consistent across generations.
What it does
Gemini Omni 1.1 Flash Reference-to-Video, in practice.
Model ID: google/gemini-omni-1.1-flash/reference-to-video Gemini Omni 1.1 Flash is Google DeepMind's natively multimodal model for video generation and editing. This variant generates a new scene conditioned on reference media — up to 10 reference images and, new in 1.1, up to 3 short reference video clips — so a character, product, or art direction stays consistent across every generation.
- Flexible resolution output — new 4K and 1080p rendering for finished work, plus a fast 360p draft mode for previewing a shot before committing to a full-quality render.
- Longer continuous generation — shots can be extended segment by segment to a total of up to 40 seconds of coherent video.
- Precise shot start and end control — first and last frame can both be supplied, producing smoother camera moves, scene transitions, and seamless loops.
- Character and style consistency — a new video reference capability (clips of up to 3 seconds) substantially improves subject and art-direction stability. This is the endpoint that exposes it.
- Subject and style consistency — carry a referenced character, object, or look across scenes and generations.
- Multi-reference conditioning — blend up to 10 reference images and up to 3 reference clips to guide subject, scene, motion, and style at once.
Run Gemini Omni 1.1 Flash Reference-to-Video
from $0.055/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | Text prompt for generation. Describes the target content, style, camera language, or character actions. Maximum 20,000 characters. You can refer to an uploaded reference by tag: <IMAGE_REF_N> is the Nth entry in reference_images and <VIDEO_ | |
reference_images | Images to use as character, scene, or style references. Accepts 1 to 10 images. Supported formats: PNG, JPEG, JPG, WebP. Each image is limited to 20MB. Supports both a public URL and a base64-encoded image for each item. | |
reference_videos | Video clips to use as motion, character, or style references. Accepts 1 to 3 clips, each up to 3 seconds long. Supported format: MP4. Each item must be a publicly accessible URL. | |
duration | The duration of the generated video in seconds. | default 10 |
aspect_ratio | The aspect ratio of the generated video. | 16:9 9:16 |
resolution | The resolution of the generated video. 360p is a fast, low-cost draft mode for previewing a shot; 1080p and 4k are upscaled from the natively generated frames. | 360p 720p 1080p 4k |
thinking_level | Controls the amount of internal reasoning the model performs before generating a response. Higher levels may improve quality on complex tasks but increase latency. | default high low |
seed | The random seed to use for the generation. -1 means a random seed will be used. | default -1 |
Sample prompt
The prompt behind the sample.
A surreal cinematic landscape at golden hour, where a colossal organic ring-shaped monument<IMAGE_REF_0> rises from an endless grassland, its massive curved surface resembling ancient weathered wood and stone. A peaceful herd of cows slowly grazes in the foreground, their movements natural and subtle as tall grass sways gently in the wind. The camera begins with a wide shot at ground level, slowly dollying forward through the grass toward the enormous ring. Warm sunlight gradually travels along the curved edge of the monument, creating a soft golden rim light. Beyond the opening of the ring,
duration: 7aspect_ratio: 9:16resolution: 720pthinking_level: defaultseed: -1FAQ
Short answers.
How much does Gemini Omni 1.1 Flash Reference-to-Video cost?
Pricing starts at $0.055 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.44. Usage is billed per request from your balance — no subscription.
Does Gemini Omni 1.1 Flash Reference-to-Video run uncensored here?
No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Gemini Omni 1.1 Flash Reference-to-Video for everything else it does well.
What does Gemini Omni 1.1 Flash Reference-to-Video take as input?
It is a text-to-video model. A natively multimodal Google DeepMind model that generates cinematic, natively sound-enabled videos from a text prompt plus up to 10 reference images and 3 reference video clips, keeping a character, product, or art direction consistent across generations.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related