Kling v3.0 Std Image-to-Video
Kling v3.0 Standard Image-to-Video model by Kuaishou. High-quality video generation from images.
What it does
Kling v3.0 Std Image-to-Video, in practice.
Kling V3.0 Standard Image-to-Video is Kuaishou's latest image-to-video generation model. Upload a reference image and describe the motion — the model generates cinematic video with optional synchronized sound, voice support, and start-to-end frame guidance.
- Latest Kling generation V3.0 delivers improved motion quality and visual fidelity over V2.6.
- Start-end frame guidance Optional end image for controlled transitions between two frames.
- Sound generation Optional synchronized sound effects generated alongside the video.
- Voice list support Add up to 2 custom voice entries for character dialogue.
- CFG scale control Fine-tune the balance between prompt adherence and creative freedom.
Run Kling v3.0 Std Image-to-Video
from $0.11/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
cfg_scale | Flexibility in video generation; The higher the value, the lower the model's degree of flexibility, and the stronger the relevance to the user's prompt. | default 0.5 |
duration | The duration of the generated media in seconds (3-15). | 3 4 5 6 7 8 9 10 11 12 13 14 |
end_image | URL of the ending image. | |
image | Supported image formats: .jpg/.jpeg/.png. The size of the image file should not exceed 10MB, the width and height of the image should be no less than 300px, and the aspect ratio of the image should be between 1:2.5 and 2.5:1. | |
negative_prompt | The negative prompt for the generation. | |
prompt | The positive prompt for the generation. Maximum 2,500 characters; longer prompts will fail Kling generation, including SR modes. | |
sound | Whether sound is generated simultaneously when generating a video. | default True |
resolution | Native 720P uses the original Kling v3.0 Std route. 1080P-SR and 1440P-SR are FlashVSR super-resolution modes generated from the native 720P source. | 720P 1080P-SR 1440P-SR |
multi_shot | Whether to enable multi-shot generation. | default |
shot_type | Multi-shot mode. customize = caller provides per-shot prompts; intelligence = model auto-splits the top-level prompt into shots. Required when multi_shot=true. | customize intelligence |
multi_prompt | Per-shot storyboards. Required when multi_shot=true and shot_type=customize. Sum of each shot's duration must equal the top-level duration; each shot duration must be >= 1. | |
elements | Subject references . Each item either references an existing subject by element_id, or creates a new one inline with element_name + reference_type + frontal_image / refer_images / refer_videos (At |
Sample prompt
The prompt behind the sample.
A minimal cube slowly moving in a dark void. Soft ambient lighting highlights its clean edges. Smooth, steady motion with a seamless loop. High contrast, ultra clean composition, 4K.
cfg_scale: 0.5duration: 5sound: FAQ
Short answers.
How much does Kling v3.0 Std Image-to-Video cost?
Pricing starts at $0.11 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.85. Usage is billed per request from your balance — no subscription.
Does Kling v3.0 Std Image-to-Video run uncensored here?
No. Kuaishou applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Kling v3.0 Std Image-to-Video for everything else it does well.
What does Kling v3.0 Std Image-to-Video take as input?
It is a image-to-video model. Kling v3.0 Standard Image-to-Video model by Kuaishou. High-quality video generation from images.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related