minimaxiH3 Get access
QwenText-to-videoVendor content policy applies

Wan-2.5 Text-to-video

A speed-optimized text-to-video option that prioritizes lower latency while retaining strong visual fidelity. Ideal for iteration, batch generation, and prompt testing.

$0.053per secondStarting price at the base resolution and quality tier.
$0.42Typical 8-second clip at the base tier. Longer or higher-resolution outputs scale with duration and tier.
Pay per useBilled per request from your balance. No subscription, no minimum.
Sample output · prompt: “A middle-aged man sitting at a wooden desk in a cozy study room, surrounded by bookshelves and a warm lamp glow. He opens an old book and reads aloud with a calm, deep voice: 'History teaches us more than just facts… it …”

What it does

Wan-2.5 Text-to-video, in practice.

Wan 2.5 is a state-of-the-art, open-source video foundation model developed by Alibaba's Wan AI team. It is designed to generate high-quality, cinematic videos complete with synchronized audio directly from text or image prompts. The model represents a significant advancement in the field of generative AI, aiming to lower the barrier for creative video production. Its core contribution lies in its ability to produce coherent, dynamic, and narratively consistent video clips with a high degree of realism and integrated audio-visual elements, such as lip-sync and sound effects, in a single, streamlined process.

  • Unified Audio-Visual Synthesis: Unlike many models that require separate steps for video and audio generation, Wan 2.5 creates video with natively synchronized audio, including voice, sound effects, and lip-sync, in one step.
  • High-Fidelity, High-Resolution Output: The model is capable of generating videos in multiple resolutions, including 480p, 720p, and full 1080p HD, with significant improvements in visual quality and frame-to-frame stability over its predecessors.
  • Extended Video Duration: Wan 2.5 can generate video clips up to 10 seconds in length, offering more creative flexibility for storytelling compared to other models in its class.
  • Advanced Cinematic Control: The model demonstrates a sophisticated understanding of cinematic language, allowing for precise control over camera movement, shot composition, and character consistency within scenes.
  • Open-Source Commitment: Following the precedent set by earlier versions, the Wan series of models, including Wan 2.5, are open-sourced to encourage research, development, and innovation within the broader AI community.
  • Content Creation: Generating short-form videos for social media, marketing campaigns, and digital advertising.

Run Wan-2.5 Text-to-video

from $0.053/sec
duration
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
audioAudio URL to guide generation (optional).
durationThe duration of the generated media in seconds.5 10
enable_prompt_expansionIf set to true, the prompt optimizer will be enabled.default
negative_promptNegative prompt for the generation.
promptThe prompt for generating the output.
seedThe random seed to use for the generation. -1 means a random seed will be used.default -1
sizeThe size of the generated media in pixels (width*height).832*480 480*832 624*624 1280*720 720*1280 960*960 1088*832 832*1088 1920*1080 1080*1920 1440*1440 1632*1248
generate_audioWhether to automatically add audio to the generated video.default True

Sample prompt

The prompt behind the sample.

A middle-aged man sitting at a wooden desk in a cozy study room, surrounded by bookshelves and a warm lamp glow. He opens an old book and reads aloud with a calm, deep voice: 'History teaches us more than just facts… it shows us who we are.' The room has subtle background sounds: pages turning, the faint ticking of a clock, and distant rain against the window.
seed: 1010059064size: 1280*720duration: 5enable_prompt_expansion:

FAQ

Short answers.

How much does Wan-2.5 Text-to-video cost?

Pricing starts at $0.053 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.42. Usage is billed per request from your balance — no subscription.

Does Wan-2.5 Text-to-video run uncensored here?

No. Qwen applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Wan-2.5 Text-to-video for everything else it does well.

What does Wan-2.5 Text-to-video take as input?

It is a text-to-video model. A speed-optimized text-to-video option that prioritizes lower latency while retaining strong visual fidelity. Ideal for iteration, batch generation, and prompt testing.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run Wan-2.5 Text-to-video.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models