minimaxiH3 Get access
QwenText-to-videoVendor content policy applies

Wan-2.5 Text-to-video Fast

Convert prompts into cinematic video clips with synchronized sound. Wan 2.5 generates 480p/720p/1080p outputs with stable motion, native audio sync, and prompt-faithful visual storytelling.

$0.11per secondStarting price at the base resolution and quality tier.
$0.85Typical 8-second clip at the base tier. Longer or higher-resolution outputs scale with duration and tier.
Pay per useBilled per request from your balance. No subscription, no minimum.
Sample output · prompt: “A confident woman with vibrant-colored braids, performing a high-energy hip-hop dance routine in the middle of a bustling intersection like Shibuya Crossing. She's wearing stylish, baggy streetwear. The crowd around her …”

What it does

Wan-2.5 Text-to-video Fast, in practice.

Wan 2.5 is a state-of-the-art, open-source video foundation model developed by Alibaba's Wan AI team. It is designed to generate high-quality, cinematic videos complete with synchronized audio directly from text or image prompts. The model represents a significant advancement in the field of generative AI, aiming to lower the barrier for creative video production. Its core contribution lies in its ability to produce coherent, dynamic, and narratively consistent video clips with a high degree of realism and integrated audio-visual elements, such as lip-sync and sound effects, in a single, streamlined process.

  • Unified Audio-Visual Synthesis: Unlike many models that require separate steps for video and audio generation, Wan 2.5 creates video with natively synchronized audio, including voice, sound effects, and lip-sync, in one step.
  • High-Fidelity, High-Resolution Output: The model is capable of generating videos in multiple resolutions, including 480p, 720p, and full 1080p HD, with significant improvements in visual quality and frame-to-frame stability over its predecessors.
  • Extended Video Duration: Wan 2.5 can generate video clips up to 10 seconds in length, offering more creative flexibility for storytelling compared to other models in its class.
  • Advanced Cinematic Control: The model demonstrates a sophisticated understanding of cinematic language, allowing for precise control over camera movement, shot composition, and character consistency within scenes.
  • Open-Source Commitment: Following the precedent set by earlier versions, the Wan series of models, including Wan 2.5, are open-sourced to encourage research, development, and innovation within the broader AI community.
  • Content Creation: Generating short-form videos for social media, marketing campaigns, and digital advertising.

Run Wan-2.5 Text-to-video Fast

from $0.11/sec
duration
size
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
audioAudio URL to guide generation (optional).
durationThe duration of the generated media in seconds.5 10
enable_prompt_expansionIf set to true, the prompt optimizer will be enabled.default
negative_promptNegative prompt for the generation.
promptThe prompt for generating the output.
seedThe random seed to use for the generation. -1 means a random seed will be used.default -1
sizeThe size of the generated media in pixels (width*height).1280*720 720*1280 1920*1080 1080*1920

Sample prompt

The prompt behind the sample.

A confident woman with vibrant-colored braids, performing a high-energy hip-hop dance routine in the middle of a bustling intersection like Shibuya Crossing. She's wearing stylish, baggy streetwear. The crowd around her is a motion blur, making her the sharp focus of the scene. Neon lights from the surrounding buildings reflect off the wet pavement. The camera uses dynamic, low-angle shots and quick cuts synchronized to the beat of a powerful trap song. Energetic, urban, rebellious.
seed: 1261673204size: 1280*720duration: 5enable_prompt_expansion:

FAQ

Short answers.

How much does Wan-2.5 Text-to-video Fast cost?

Pricing starts at $0.11 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.85. Usage is billed per request from your balance — no subscription.

Does Wan-2.5 Text-to-video Fast run uncensored here?

No. Qwen applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Wan-2.5 Text-to-video Fast for everything else it does well.

What does Wan-2.5 Text-to-video Fast take as input?

It is a text-to-video model. Convert prompts into cinematic video clips with synchronized sound. Wan 2.5 generates 480p/720p/1080p outputs with stable motion, native audio sync, and prompt-faithful visual storytelling.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run Wan-2.5 Text-to-video Fast.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models