minimaxiH3 Get access
GoogleText-to-videoVendor content policy applies

Gemini Omni 1.1 Flash Reference-to-Video

A natively multimodal Google DeepMind model that generates cinematic, natively sound-enabled videos from a text prompt plus up to 10 reference images and 3 reference video clips, keeping a character, product, or art direction consistent across generations.

$0.055per secondStarting price at the base resolution and quality tier.
$0.44Typical 8-second clip at the base tier. Longer or higher-resolution outputs scale with duration and tier.
Pay per useBilled per request from your balance. No subscription, no minimum.
Sample output · prompt: “A surreal cinematic landscape at golden hour, where a colossal organic ring-shaped monument<IMAGE_REF_0> rises from an endless grassland, its massive curved surface resembling ancient weathered wood and stone. A peaceful…”

What it does

Gemini Omni 1.1 Flash Reference-to-Video, in practice.

Model ID: google/gemini-omni-1.1-flash/reference-to-video Gemini Omni 1.1 Flash is Google DeepMind's natively multimodal model for video generation and editing. This variant generates a new scene conditioned on reference media — up to 10 reference images and, new in 1.1, up to 3 short reference video clips — so a character, product, or art direction stays consistent across every generation.

  • Flexible resolution output — new 4K and 1080p rendering for finished work, plus a fast 360p draft mode for previewing a shot before committing to a full-quality render.
  • Longer continuous generation — shots can be extended segment by segment to a total of up to 40 seconds of coherent video.
  • Precise shot start and end control — first and last frame can both be supplied, producing smoother camera moves, scene transitions, and seamless loops.
  • Character and style consistency — a new video reference capability (clips of up to 3 seconds) substantially improves subject and art-direction stability. This is the endpoint that exposes it.
  • Subject and style consistency — carry a referenced character, object, or look across scenes and generations.
  • Multi-reference conditioning — blend up to 10 reference images and up to 3 reference clips to guide subject, scene, motion, and style at once.

Run Gemini Omni 1.1 Flash Reference-to-Video

from $0.055/sec
aspect_ratio
resolution
thinking_level
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
promptText prompt for generation. Describes the target content, style, camera language, or character actions. Maximum 20,000 characters. You can refer to an uploaded reference by tag: <IMAGE_REF_N> is the Nth entry in reference_images and <VIDEO_
reference_imagesImages to use as character, scene, or style references. Accepts 1 to 10 images. Supported formats: PNG, JPEG, JPG, WebP. Each image is limited to 20MB. Supports both a public URL and a base64-encoded image for each item.
reference_videosVideo clips to use as motion, character, or style references. Accepts 1 to 3 clips, each up to 3 seconds long. Supported format: MP4. Each item must be a publicly accessible URL.
durationThe duration of the generated video in seconds.default 10
aspect_ratioThe aspect ratio of the generated video.16:9 9:16
resolutionThe resolution of the generated video. 360p is a fast, low-cost draft mode for previewing a shot; 1080p and 4k are upscaled from the natively generated frames.360p 720p 1080p 4k
thinking_levelControls the amount of internal reasoning the model performs before generating a response. Higher levels may improve quality on complex tasks but increase latency.default high low
seedThe random seed to use for the generation. -1 means a random seed will be used.default -1

Sample prompt

The prompt behind the sample.

A surreal cinematic landscape at golden hour, where a colossal organic ring-shaped monument<IMAGE_REF_0> rises from an endless grassland, its massive curved surface resembling ancient weathered wood and stone. A peaceful herd of cows slowly grazes in the foreground, their movements natural and subtle as tall grass sways gently in the wind. The camera begins with a wide shot at ground level, slowly dollying forward through the grass toward the enormous ring. Warm sunlight gradually travels along the curved edge of the monument, creating a soft golden rim light. Beyond the opening of the ring,
duration: 7aspect_ratio: 9:16resolution: 720pthinking_level: defaultseed: -1

FAQ

Short answers.

How much does Gemini Omni 1.1 Flash Reference-to-Video cost?

Pricing starts at $0.055 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.44. Usage is billed per request from your balance — no subscription.

Does Gemini Omni 1.1 Flash Reference-to-Video run uncensored here?

No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Gemini Omni 1.1 Flash Reference-to-Video for everything else it does well.

What does Gemini Omni 1.1 Flash Reference-to-Video take as input?

It is a text-to-video model. A natively multimodal Google DeepMind model that generates cinematic, natively sound-enabled videos from a text prompt plus up to 10 reference images and 3 reference video clips, keeping a character, product, or art direction consistent across generations.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run Gemini Omni 1.1 Flash Reference-to-Video.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models