Gemini Omni Flash Reference-to-Video Developer
Gemini Omni Flash is Google's multimodal video generation model. This reference-to-video variant transforms existing video clips using reference images and text prompts, enabling video style transfer, scene editing, and character insertion.
What it does
Gemini Omni Flash Reference-to-Video Developer, in practice.
Model ID: google/gemini-omni-flash/reference-to-video-developer Gemini Omni is Google's multimodal video generation model designed to create high-quality video content from diverse input types. This variant accepts a text prompt, reference images, and a source video clip, enabling the most expressive form of video generation: transforming existing footage while preserving coherence and injecting new creative direction.
- Video-guided generation — Provide a source video clip as a structural or stylistic reference; the model builds upon it to produce new content.
- Image reference anchoring — Supply 1 to 5 reference images alongside the video to define subjects, characters, or visual style.
- Precise clip trimming — Specify `start` and `end` timestamps within the source video to use only the most relevant segment (trim window ≤ 10 seconds).
- Rich prompt understanding — Describe transformations, additions, camera language, and mood in a prompt of up to 20,000 characters.
- Multi-resolution output — Generate at 720p, 1080p, or 4K.
- Flexible aspect ratios — 16:9 landscape or 9:16 portrait.
Run Gemini Omni Flash Reference-to-Video Developer
from $0.18/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | Text prompt for generation. Describes the target content, style, camera language, or character actions. Maximum 20,000 characters. | |
images | Images to use as character, scene, or style references. Accepts 1 to 5 images when combined with a video reference (video costs 2 units out of a total quota of 7). Supported formats: PNG, JPEG, JPG, WebP. Each image is limited to 20MB. | |
video_clips | Source video clips to use as references for generation. Supports 1 video clip. Each video is limited to 100MB and 30 seconds duration. The trimmed segment (ends - start) must not exceed 10 seconds. | |
duration | The duration of the generated video in seconds. | 4 6 8 10 |
aspect_ratio | The aspect ratio of the generated video. | 16:9 9:16 |
resolution | The resolution of the generated video. | 720p 1080p 4k |
seed | Random seed for reproducibility. Use -1 to use a random seed. | default -1 |
Sample prompt
The prompt behind the sample.
replace the bee in the video with the butterfly in the first image
seed: -1duration: 10resolution: 1080paspect_ratio: 16:9FAQ
Short answers.
How much does Gemini Omni Flash Reference-to-Video Developer cost?
Pricing starts at $0.18 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $1.44. Usage is billed per request from your balance — no subscription.
Does Gemini Omni Flash Reference-to-Video Developer run uncensored here?
No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Gemini Omni Flash Reference-to-Video Developer for everything else it does well.
What does Gemini Omni Flash Reference-to-Video Developer take as input?
It is a video-to-video model. Gemini Omni Flash is Google's multimodal video generation model. This reference-to-video variant transforms existing video clips using reference images and text prompts, enabling video style transfer, scene editing, and character insertion.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related