minimaxiH3 Get access
GoogleImage-to-videoVendor content policy applies

Gemini Omni Flash Image-to-Video Developer

Gemini Omni Flash is Google's multimodal video generation model. This image-to-video variant creates subject-consistent videos from up to 7 reference images combined with a text prompt, preserving visual identity across the full generated video.

$0.17per secondStarting price at the base resolution and quality tier.
$1.34Typical 8-second clip at the base tier. Longer or higher-resolution outputs scale with duration and tier.
Pay per useBilled per request from your balance. No subscription, no minimum.
Sample output · prompt: “Style & Camera: High-speed cinematic action photography, 4K resolution, shallow depth of field. The camera begins with a low-angle tracking shot, moving parallel to a soccer player sprinting down a rain-slicked green pit…”

What it does

Gemini Omni Flash Image-to-Video Developer, in practice.

Model ID: google/gemini-omni-flash/image-to-video-developer Gemini Omni is Google's multimodal video generation model designed to create high-quality video content from diverse input types. This variant accepts a text prompt plus up to 7 reference images, enabling subject-consistent video generation where the visual identity of characters, objects, or scenes is anchored by real image references.

  • Image-guided generation — Provide 1 to 7 reference images to anchor subjects, environments, or visual styles.
  • Subject consistency — The model preserves key visual details from the reference images across the full video duration.
  • Rich prompt understanding — Complement image references with a prompt of up to 20,000 characters describing actions, camera movements, lighting, and mood.
  • Multi-resolution output — Generate at 720p, 1080p, or 4K.
  • Flexible aspect ratios — 16:9 landscape or 9:16 portrait.
  • Controllable duration — 4, 6, 8, or 10 seconds per generation.

Run Gemini Omni Flash Image-to-Video Developer

from $0.17/sec
Drop an image hereor click to choose a fileFiles stay in your project. Nothing is trained on.
duration
aspect_ratio
resolution
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
promptText prompt for generation. Describes the target content, style, camera language, or character actions. Maximum 20,000 characters.
imagesImages to use as character, scene, or style references. Accepts 1 to 7 images. Supported formats: PNG, JPEG, JPG, WebP. Each image is limited to 20MB.
durationThe duration of the generated video in seconds.4 6 8 10
aspect_ratioThe aspect ratio of the generated video.16:9 9:16
resolutionThe resolution of the generated video.720p 1080p 4k
seedRandom seed for reproducibility. Use -1 to use a random seed.default -1

Sample prompt

The prompt behind the sample.

Style & Camera: High-speed cinematic action photography, 4K resolution, shallow depth of field. The camera begins with a low-angle tracking shot, moving parallel to a soccer player sprinting down a rain-slicked green pitch under dramatic stadium floodlights. Visual & Action Sequence: The soccer player, wearing a detailed jersey with realistic fabric ripples, powerful plants their left foot into the wet turf. Water droplets and blades of grass explode upwards from the impact. As they swing their right leg back and powerfully strike the soccer ball, the ball deforms slightly against the boot be
duration: 6aspect_ratio: 16:9resolution: 720pseed: -1

FAQ

Short answers.

How much does Gemini Omni Flash Image-to-Video Developer cost?

Pricing starts at $0.17 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $1.34. Usage is billed per request from your balance — no subscription.

Does Gemini Omni Flash Image-to-Video Developer run uncensored here?

No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Gemini Omni Flash Image-to-Video Developer for everything else it does well.

What does Gemini Omni Flash Image-to-Video Developer take as input?

It is a image-to-video model. Gemini Omni Flash is Google's multimodal video generation model. This image-to-video variant creates subject-consistent videos from up to 7 reference images combined with a text prompt, preserving visual identity across the full generated video.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run Gemini Omni Flash Image-to-Video Developer.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models