minimaxiH3 Get access
KuaishouText-to-videoVendor content policy applies

Kling Video O1 Text-to-video

Kling Omni Video O1 is Kuaishou's first unified multi-modal video model with MVL (Multi-modal Visual Language) technology. Text-to-Video mode generates cinematic videos from text prompts with subject consistency, natural physics simulation, and precise semantic understanding. Ready-to-use REST API, best performance, no coldstarts, affordable pricing.

$0.14per secondStarting price at the base resolution and quality tier.
$1.14Typical 8-second clip at the base tier. Longer or higher-resolution outputs scale with duration and tier.
Pay per useBilled per request from your balance. No subscription, no minimum.
Sample output · prompt: “Cyberpunk action scene. Two cyborg samurais dash past each other on a rainy street. They strike with glowing laser katanas. A burst of blue and red sparks explodes in the center of the frame. The camera pans quickly to f…”

What it does

Kling Video O1 Text-to-video, in practice.

Kling Omni Video O1 is Kuaishou's groundbreaking unified multi-modal video model, representing the world's first AI system that seamlessly integrates text, images, videos, and subject references into a single creative engine. The Text-to-Video mode transforms natural language prompts into stunning, cinematic video content.

  • Text-to-video generation
  • Image-to-video transformation
  • Reference-based video creation
  • Video editing and modification
  • Shot extension and scene continuation
  • Natural language descriptions

Run Kling Video O1 Text-to-video

from $0.14/sec
aspect_ratio
duration
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
aspect_ratioThe aspect ratio of the generated video.16:9 9:16 1:1
durationThe duration of the generated media in seconds.5 10
promptThe positive prompt for the generation.

Sample prompt

The prompt behind the sample.

Cyberpunk action scene. Two cyborg samurais dash past each other on a rainy street. They strike with glowing laser katanas. A burst of blue and red sparks explodes in the center of the frame. The camera pans quickly to follow the movement. Unreal Engine 5 render, high octane.
duration: 5aspect_ratio: 16:9

FAQ

Short answers.

How much does Kling Video O1 Text-to-video cost?

Pricing starts at $0.14 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $1.14. Usage is billed per request from your balance — no subscription.

Does Kling Video O1 Text-to-video run uncensored here?

No. Kuaishou applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Kling Video O1 Text-to-video for everything else it does well.

What does Kling Video O1 Text-to-video take as input?

It is a text-to-video model. Kling Omni Video O1 is Kuaishou's first unified multi-modal video model with MVL (Multi-modal Visual Language) technology. Text-to-Video mode generates cinematic videos from text prompts with subject consistency, natural physics simulation, and precise semanti

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run Kling Video O1 Text-to-video.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models