minimaxiH3 Get access
KuaishouImage-to-videoVendor content policy applies

Kling v3.0 Std Image-to-Video

Kling v3.0 Standard Image-to-Video model by Kuaishou. High-quality video generation from images.

$0.11per secondStarting price at the base resolution and quality tier.
$0.85Typical 8-second clip at the base tier. Longer or higher-resolution outputs scale with duration and tier.
Pay per useBilled per request from your balance. No subscription, no minimum.
Sample output · prompt: “A minimal cube slowly moving in a dark void. Soft ambient lighting highlights its clean edges. Smooth, steady motion with a seamless loop. High contrast, ultra clean composition, 4K.”

What it does

Kling v3.0 Std Image-to-Video, in practice.

Kling V3.0 Standard Image-to-Video is Kuaishou's latest image-to-video generation model. Upload a reference image and describe the motion — the model generates cinematic video with optional synchronized sound, voice support, and start-to-end frame guidance.

  • Latest Kling generation V3.0 delivers improved motion quality and visual fidelity over V2.6.
  • Start-end frame guidance Optional end image for controlled transitions between two frames.
  • Sound generation Optional synchronized sound effects generated alongside the video.
  • Voice list support Add up to 2 custom voice entries for character dialogue.
  • CFG scale control Fine-tune the balance between prompt adherence and creative freedom.

Run Kling v3.0 Std Image-to-Video

from $0.11/sec
Drop an image hereor click to choose a fileFiles stay in your project. Nothing is trained on.
resolution
shot_type
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
cfg_scaleFlexibility in video generation; The higher the value, the lower the model's degree of flexibility, and the stronger the relevance to the user's prompt.default 0.5
durationThe duration of the generated media in seconds (3-15).3 4 5 6 7 8 9 10 11 12 13 14
end_imageURL of the ending image.
imageSupported image formats: .jpg/.jpeg/.png. The size of the image file should not exceed 10MB, the width and height of the image should be no less than 300px, and the aspect ratio of the image should be between 1:2.5 and 2.5:1.
negative_promptThe negative prompt for the generation.
promptThe positive prompt for the generation. Maximum 2,500 characters; longer prompts will fail Kling generation, including SR modes.
soundWhether sound is generated simultaneously when generating a video.default True
resolutionNative 720P uses the original Kling v3.0 Std route. 1080P-SR and 1440P-SR are FlashVSR super-resolution modes generated from the native 720P source.720P 1080P-SR 1440P-SR
multi_shotWhether to enable multi-shot generation.default
shot_typeMulti-shot mode. customize = caller provides per-shot prompts; intelligence = model auto-splits the top-level prompt into shots. Required when multi_shot=true.customize intelligence
multi_promptPer-shot storyboards. Required when multi_shot=true and shot_type=customize. Sum of each shot's duration must equal the top-level duration; each shot duration must be >= 1.
elementsSubject references . Each item either references an existing subject by element_id, or creates a new one inline with element_name + reference_type + frontal_image / refer_images / refer_videos (At

Sample prompt

The prompt behind the sample.

A minimal cube slowly moving in a dark void. Soft ambient lighting highlights its clean edges. Smooth, steady motion with a seamless loop. High contrast, ultra clean composition, 4K.
cfg_scale: 0.5duration: 5sound:

FAQ

Short answers.

How much does Kling v3.0 Std Image-to-Video cost?

Pricing starts at $0.11 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.85. Usage is billed per request from your balance — no subscription.

Does Kling v3.0 Std Image-to-Video run uncensored here?

No. Kuaishou applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Kling v3.0 Std Image-to-Video for everything else it does well.

What does Kling v3.0 Std Image-to-Video take as input?

It is a image-to-video model. Kling v3.0 Standard Image-to-Video model by Kuaishou. High-quality video generation from images.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run Kling v3.0 Std Image-to-Video.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models