Nano Banana 2 Lite Reference-to-image
Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image) is Google's fastest, most cost-efficient image model, turning a source video clip plus a natural-language prompt (and optionally up to 14 reference images) into brand-new still images at low latency.

What it does
Nano Banana 2 Lite Reference-to-image, in practice.
Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image, gemini-3.1-flash-lite-image) is Google's fastest and most cost-efficient image model in the Nano Banana family. This variant is driven by a source video clip plus a natural-language prompt: the model uses the video's context as a multimodal reference, extracts its visual themes and key moments, and synthesizes brand-new still images from them — all with the low latency that defines the Lite tier. It shares the underlying model with the text-to-image and edit variants — the difference is the input. Instead of a text prompt alone or a set of reference photos, you supply one video clip (and, optionally, up to 14 reference images) and describe the
- Video as a visual reference — Turn footage into stills. The model analyzes video frames to understand subjects, scenes, and key events, then generates images grounded in that context. See Google's [video-to-image generation guide](https://ai.google.dev/gemini-api/docs/image-generation#video-to-image).
- Flexible source input — Accepts a public YouTube URL or a direct HTTP video URL (up to 15 MB), with adjustable trim range and sampling FPS so you control which part of the clip is used.
- Multimodal composition — Combine the video reference with up to 14 additional input images to blend subjects, transfer styles, or assemble scenes from multiple sources.
- Strong character consistency — Preserves character identities and object fidelity drawn from the source video, so subjects stay recognizable.
- Legible in-image text — Renders and localizes readable text directly within generated images for quick captioning and design work.
- World knowledge — Understands scene structure and real-world context to keep generated frames coherent and plausible.
Run Nano Banana 2 Lite Reference-to-image
from $0.060/imageParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
aspect_ratio | The aspect ratio of the generated media. | auto 1:1 3:2 2:3 3:4 4:3 4:5 5:4 9:16 16:9 21:9 4:1 |
images | List of URLs of input images for editing. The maximum number of images is 14. | |
prompt | The positive prompt for the generation. | |
resolution | The resolution of the output image. | 1k |
thinking_level | Controls the amount of internal reasoning the model performs before generating a response. Higher levels may improve quality on complex tasks but increase latency. | default high minimal |
video_clips | Source video clips to use as references for generation. Supports 1 video clip. | |
seed | Random seed. Does not guarantee determinism but may improve repeatability. -1 means a random seed will be used. | default -1 |
top_p | Probability threshold for top-p sampling | default 0.95 |
temperature | Creativity allowed in the responses. Best results at default 1.0. Lower values may impact reasoning. | default 1 |
Sample prompt
The prompt behind the sample.
The car sped along the road.
aspect_ratio: 16:9enable_base64_output: enable_sync_mode: resolution: 1kthinking_level: defaultFAQ
Short answers.
How much does Nano Banana 2 Lite Reference-to-image cost?
Pricing starts at $0.060 per image at the base resolution; higher resolutions and longer durations cost more. Usage is billed per request from your balance — no subscription.
Does Nano Banana 2 Lite Reference-to-image run uncensored here?
No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Nano Banana 2 Lite Reference-to-image for everything else it does well.
What does Nano Banana 2 Lite Reference-to-image take as input?
It is a image editing model. Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image) is Google's fastest, most cost-efficient image model, turning a source video clip plus a natural-language prompt (and optionally up to 14 reference images) into brand-new still images at low latency.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related