minimaxiH3 Get access
CommunityImage-to-videoVendor content policy applies

Nvidia Cosmos 3 Super Image-to-Video

Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs.

$0.083per secondStarting price at the base resolution and quality tier.
$0.66Typical 8-second clip at the base tier. Longer or higher-resolution outputs scale with duration and tier.
Pay per useBilled per request from your balance. No subscription, no minimum.
Sample output · prompt: “A cute 3D cartoon girl wearing fluffy earmuffs and oversized headphones smiles brightly at the camera. She gently raises her index finger as if sharing an exciting idea, blinking naturally with expressive eyes. Her hair …”

What it does

Nvidia Cosmos 3 Super Image-to-Video, in practice.

NVIDIA Cosmos 3 Super Image-to-Video is NVIDIA's flagship image animation model built on the Cosmos 3 architecture. It transforms a single still image into a smooth, physically grounded video clip guided by a text prompt — delivering cinematic motion, accurate temporal coherence, and fine-grained control over frame count and pacing.

Run Nvidia Cosmos 3 Super Image-to-Video

from $0.083/sec
Drop an image hereor click to choose a fileFiles stay in your project. Nothing is trained on.
duration
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
promptText description of the motion and scene to generate. Maximum 2,500 characters.
image_urlThe reference image to animate. Accepts a public URL or Base64-encoded image. Supported formats: png, jpeg, jpg, webp.
negative_promptText describing content to exclude from the generated video.
image_sizeOutput image resolution. Use a preset size or specify custom width and height via an ImageSize object.default landscape_16_9
durationThe duration of the generated media in seconds (1-7).1 2 3 4 5 6 7
num_inference_stepsNumber of denoising steps. Higher values improve quality but increase generation time.default 28
guidance_scaleClassifier-free guidance scale. Higher values make the output follow the prompt more closely.default 6.0
seedRandom seed for reproducibility. Use the same seed and prompt to reproduce results.
enable_safety_checkerEnable or disable the safety checker. Disabling the safety checker may result in NSFW content.default

Sample prompt

The prompt behind the sample.

A cute 3D cartoon girl wearing fluffy earmuffs and oversized headphones smiles brightly at the camera. She gently raises her index finger as if sharing an exciting idea, blinking naturally with expressive eyes. Her hair softly sways as a gentle breeze passes by. The camera begins with a medium shot, then slowly pushes into a close-up, capturing her warm smile and sparkling eyes. Soft studio lighting creates delicate shadows while floating dust particles shimmer in the air. The background remains clean and minimalistic with warm brown tones. Pixar-quality animation, cinematic lighting, subtle f
negative_prompt: image_size: landscape_16_9duration: 5num_inference_steps: 28guidance_scale: 6

FAQ

Short answers.

How much does Nvidia Cosmos 3 Super Image-to-Video cost?

Pricing starts at $0.083 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.66. Usage is billed per request from your balance — no subscription.

Does Nvidia Cosmos 3 Super Image-to-Video run uncensored here?

No. Community applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Nvidia Cosmos 3 Super Image-to-Video for everything else it does well.

What does Nvidia Cosmos 3 Super Image-to-Video take as input?

It is a image-to-video model. Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run Nvidia Cosmos 3 Super Image-to-Video.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models