minimaxiH3 Get access
GoogleVideo-to-videoVendor content policy applies

Gemini Omni Flash Reference-to-Video Developer

Gemini Omni Flash is Google's multimodal video generation model. This reference-to-video variant transforms existing video clips using reference images and text prompts, enabling video style transfer, scene editing, and character insertion.

$0.18per secondStarting price at the base resolution and quality tier.
$1.44Typical 8-second clip at the base tier. Longer or higher-resolution outputs scale with duration and tier.
Pay per useBilled per request from your balance. No subscription, no minimum.
Sample output · prompt: “replace the bee in the video with the butterfly in the first image”

What it does

Gemini Omni Flash Reference-to-Video Developer, in practice.

Model ID: google/gemini-omni-flash/reference-to-video-developer Gemini Omni is Google's multimodal video generation model designed to create high-quality video content from diverse input types. This variant accepts a text prompt, reference images, and a source video clip, enabling the most expressive form of video generation: transforming existing footage while preserving coherence and injecting new creative direction.

  • Video-guided generation — Provide a source video clip as a structural or stylistic reference; the model builds upon it to produce new content.
  • Image reference anchoring — Supply 1 to 5 reference images alongside the video to define subjects, characters, or visual style.
  • Precise clip trimming — Specify `start` and `end` timestamps within the source video to use only the most relevant segment (trim window ≤ 10 seconds).
  • Rich prompt understanding — Describe transformations, additions, camera language, and mood in a prompt of up to 20,000 characters.
  • Multi-resolution output — Generate at 720p, 1080p, or 4K.
  • Flexible aspect ratios — 16:9 landscape or 9:16 portrait.

Run Gemini Omni Flash Reference-to-Video Developer

from $0.18/sec
Drop a video hereor click to choose a fileFiles stay in your project. Nothing is trained on.
duration
aspect_ratio
resolution
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
promptText prompt for generation. Describes the target content, style, camera language, or character actions. Maximum 20,000 characters.
imagesImages to use as character, scene, or style references. Accepts 1 to 5 images when combined with a video reference (video costs 2 units out of a total quota of 7). Supported formats: PNG, JPEG, JPG, WebP. Each image is limited to 20MB.
video_clipsSource video clips to use as references for generation. Supports 1 video clip. Each video is limited to 100MB and 30 seconds duration. The trimmed segment (ends - start) must not exceed 10 seconds.
durationThe duration of the generated video in seconds.4 6 8 10
aspect_ratioThe aspect ratio of the generated video.16:9 9:16
resolutionThe resolution of the generated video.720p 1080p 4k
seedRandom seed for reproducibility. Use -1 to use a random seed.default -1

Sample prompt

The prompt behind the sample.

replace the bee in the video with the butterfly in the first image
seed: -1duration: 10resolution: 1080paspect_ratio: 16:9

FAQ

Short answers.

How much does Gemini Omni Flash Reference-to-Video Developer cost?

Pricing starts at $0.18 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $1.44. Usage is billed per request from your balance — no subscription.

Does Gemini Omni Flash Reference-to-Video Developer run uncensored here?

No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Gemini Omni Flash Reference-to-Video Developer for everything else it does well.

What does Gemini Omni Flash Reference-to-Video Developer take as input?

It is a video-to-video model. Gemini Omni Flash is Google's multimodal video generation model. This reference-to-video variant transforms existing video clips using reference images and text prompts, enabling video style transfer, scene editing, and character insertion.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run Gemini Omni Flash Reference-to-Video Developer.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models