minimaxiH3 Get access
CommunityText-to-imageVendor content policy applies

Nvidia Cosmos 3 Super Text-to-Image

Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs.

$0.066per imageStarting price at the base resolution and quality tier.
$0.66Ten outputs at the base tier. Larger sizes and premium quality tiers cost more.
Pay per useBilled per request from your balance. No subscription, no minimum.
Nvidia Cosmos 3 Super Text-to-Image sample output
Sample output · prompt: “A symphony orchestra performing inside a grand European train station after the last train has departed. Empty platforms stretch into darkness while warm lantern light illuminates the musicians. Steam drifts through the …”

What it does

Nvidia Cosmos 3 Super Text-to-Image, in practice.

NVIDIA Cosmos 3 Super Text-to-Image is NVIDIA's flagship text-to-image model built on the Cosmos 3 architecture. Designed for high-fidelity physical world simulation, it delivers photorealistic images with strong prompt adherence, precise spatial reasoning, and support for negative prompting — making it well suited for both creative and technical generation tasks.

Run Nvidia Cosmos 3 Super Text-to-Image

from $0.066/image
output_format
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
promptThe text prompt describing the image to generate. Maximum 2,500 characters.
negative_promptText describing content to exclude from the generated image.
image_sizeOutput image resolution. Use a preset size or specify custom width and height via an ImageSize object.default landscape_16_9
num_inference_stepsNumber of denoising steps. Higher values improve quality but increase generation time.default 28
guidance_scaleClassifier-free guidance scale. Higher values make the output follow the prompt more closely.default 4.0
output_formatOutput image format.jpeg png
seedRandom seed for reproducibility. Use the same seed and prompt to reproduce results.
enable_safety_checkerEnable or disable the safety checker. Disabling the safety checker may result in NSFW content.default

Sample prompt

The prompt behind the sample.

A symphony orchestra performing inside a grand European train station after the last train has departed. Empty platforms stretch into darkness while warm lantern light illuminates the musicians. Steam drifts through the air, creating a dreamlike cinematic atmosphere. Emotional storytelling, movie still, ultra realistic.
negative_prompt: image_size: landscape_16_9num_inference_steps: 28guidance_scale: 4output_format: jpeg

FAQ

Short answers.

How much does Nvidia Cosmos 3 Super Text-to-Image cost?

Pricing starts at $0.066 per image at the base resolution; higher resolutions and longer durations cost more. Usage is billed per request from your balance — no subscription.

Does Nvidia Cosmos 3 Super Text-to-Image run uncensored here?

No. Community applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Nvidia Cosmos 3 Super Text-to-Image for everything else it does well.

What does Nvidia Cosmos 3 Super Text-to-Image take as input?

It is a text-to-image model. Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run Nvidia Cosmos 3 Super Text-to-Image.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models