Kling Video O1 Text-to-video
Kling Omni Video O1 is Kuaishou's first unified multi-modal video model with MVL (Multi-modal Visual Language) technology. Text-to-Video mode generates cinematic videos from text prompts with subject consistency, natural physics simulation, and precise semantic understanding. Ready-to-use REST API, best performance, no coldstarts, affordable pricing.
What it does
Kling Video O1 Text-to-video, in practice.
Kling Omni Video O1 is Kuaishou's groundbreaking unified multi-modal video model, representing the world's first AI system that seamlessly integrates text, images, videos, and subject references into a single creative engine. The Text-to-Video mode transforms natural language prompts into stunning, cinematic video content.
- Text-to-video generation
- Image-to-video transformation
- Reference-based video creation
- Video editing and modification
- Shot extension and scene continuation
- Natural language descriptions
Run Kling Video O1 Text-to-video
from $0.14/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
aspect_ratio | The aspect ratio of the generated video. | 16:9 9:16 1:1 |
duration | The duration of the generated media in seconds. | 5 10 |
prompt | The positive prompt for the generation. |
Sample prompt
The prompt behind the sample.
Cyberpunk action scene. Two cyborg samurais dash past each other on a rainy street. They strike with glowing laser katanas. A burst of blue and red sparks explodes in the center of the frame. The camera pans quickly to follow the movement. Unreal Engine 5 render, high octane.
duration: 5aspect_ratio: 16:9FAQ
Short answers.
How much does Kling Video O1 Text-to-video cost?
Pricing starts at $0.14 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $1.14. Usage is billed per request from your balance — no subscription.
Does Kling Video O1 Text-to-video run uncensored here?
No. Kuaishou applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Kling Video O1 Text-to-video for everything else it does well.
What does Kling Video O1 Text-to-video take as input?
It is a text-to-video model. Kling Omni Video O1 is Kuaishou's first unified multi-modal video model with MVL (Multi-modal Visual Language) technology. Text-to-Video mode generates cinematic videos from text prompts with subject consistency, natural physics simulation, and precise semanti
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related