minimaxiH3 Get access
GoogleText-to-videoVendor content policy applies

Veo3.1 Text-to-video

Generate high-fidelity videos from text prompts with Google’s most advanced generative video model. Veo 3.1 delivers cinematic quality, dynamic camera motion, and lifelike detail for storytelling and creative production.

$0.30per secondStarting price at the base resolution and quality tier.
$2.40Typical 8-second clip at the base tier. Longer or higher-resolution outputs scale with duration and tier.
Pay per useBilled per request from your balance. No subscription, no minimum.
Sample output · prompt: “A car slowly driving down a quiet road, surrounded by a calm landscape, subtle motion, soft sunlight illuminating the scene, cinematic camera movement, peaceful atmosphere.”

What it does

Veo3.1 Text-to-video, in practice.

Veo 3.1 T2V is the latest text-to-video model from Google DeepMind, designed to bring cinematic storytelling to life through text. It generates high-fidelity 1080p videos with synchronized, context-aware audio, realistic motion, and narrative consistency — making it one of the most advanced generative video systems ever released.

  • Cinematic Realism
  • Native Audio Generation
  • Dialogue & Lip-Sync
  • Subject Consistency (R2V)
  • Video Interpolation
  • Flexible Output

Run Veo3.1 Text-to-video

from $0.30/sec
aspect_ratio
duration
resolution
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
aspect_ratioAspect ratio of the video.16:9 9:16
durationThe duration of the generated media in seconds.8 4 6
generate_audioWhether to generate audio.default
negative_promptNegative prompt for the generation.
promptText prompt for generation; Positive text prompt.
resolutionVideo resolution.720p 1080p 4k
seedThe random seed to use for the generation.

Sample prompt

The prompt behind the sample.

A car slowly driving down a quiet road, surrounded by a calm landscape, subtle motion, soft sunlight illuminating the scene, cinematic camera movement, peaceful atmosphere.
duration: 8resolution: 720paspect_ratio: 16:9generate_audio: True

FAQ

Short answers.

How much does Veo3.1 Text-to-video cost?

Pricing starts at $0.30 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $2.40. Usage is billed per request from your balance — no subscription.

Does Veo3.1 Text-to-video run uncensored here?

No. Google applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Veo3.1 Text-to-video for everything else it does well.

What does Veo3.1 Text-to-video take as input?

It is a text-to-video model. Generate high-fidelity videos from text prompts with Google’s most advanced generative video model. Veo 3.1 delivers cinematic quality, dynamic camera motion, and lifelike detail for storytelling and creative production.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run Veo3.1 Text-to-video.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models