Wan-2.5 Text-to-video
A speed-optimized text-to-video option that prioritizes lower latency while retaining strong visual fidelity. Ideal for iteration, batch generation, and prompt testing.
What it does
Wan-2.5 Text-to-video, in practice.
Wan 2.5 is a state-of-the-art, open-source video foundation model developed by Alibaba's Wan AI team. It is designed to generate high-quality, cinematic videos complete with synchronized audio directly from text or image prompts. The model represents a significant advancement in the field of generative AI, aiming to lower the barrier for creative video production. Its core contribution lies in its ability to produce coherent, dynamic, and narratively consistent video clips with a high degree of realism and integrated audio-visual elements, such as lip-sync and sound effects, in a single, streamlined process.
- Unified Audio-Visual Synthesis: Unlike many models that require separate steps for video and audio generation, Wan 2.5 creates video with natively synchronized audio, including voice, sound effects, and lip-sync, in one step.
- High-Fidelity, High-Resolution Output: The model is capable of generating videos in multiple resolutions, including 480p, 720p, and full 1080p HD, with significant improvements in visual quality and frame-to-frame stability over its predecessors.
- Extended Video Duration: Wan 2.5 can generate video clips up to 10 seconds in length, offering more creative flexibility for storytelling compared to other models in its class.
- Advanced Cinematic Control: The model demonstrates a sophisticated understanding of cinematic language, allowing for precise control over camera movement, shot composition, and character consistency within scenes.
- Open-Source Commitment: Following the precedent set by earlier versions, the Wan series of models, including Wan 2.5, are open-sourced to encourage research, development, and innovation within the broader AI community.
- Content Creation: Generating short-form videos for social media, marketing campaigns, and digital advertising.
Run Wan-2.5 Text-to-video
from $0.053/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
audio | Audio URL to guide generation (optional). | |
duration | The duration of the generated media in seconds. | 5 10 |
enable_prompt_expansion | If set to true, the prompt optimizer will be enabled. | default |
negative_prompt | Negative prompt for the generation. | |
prompt | The prompt for generating the output. | |
seed | The random seed to use for the generation. -1 means a random seed will be used. | default -1 |
size | The size of the generated media in pixels (width*height). | 832*480 480*832 624*624 1280*720 720*1280 960*960 1088*832 832*1088 1920*1080 1080*1920 1440*1440 1632*1248 |
generate_audio | Whether to automatically add audio to the generated video. | default True |
Sample prompt
The prompt behind the sample.
A middle-aged man sitting at a wooden desk in a cozy study room, surrounded by bookshelves and a warm lamp glow. He opens an old book and reads aloud with a calm, deep voice: 'History teaches us more than just facts… it shows us who we are.' The room has subtle background sounds: pages turning, the faint ticking of a clock, and distant rain against the window.
seed: 1010059064size: 1280*720duration: 5enable_prompt_expansion: FAQ
Short answers.
How much does Wan-2.5 Text-to-video cost?
Pricing starts at $0.053 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.42. Usage is billed per request from your balance — no subscription.
Does Wan-2.5 Text-to-video run uncensored here?
No. Qwen applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Wan-2.5 Text-to-video for everything else it does well.
What does Wan-2.5 Text-to-video take as input?
It is a text-to-video model. A speed-optimized text-to-video option that prioritizes lower latency while retaining strong visual fidelity. Ideal for iteration, batch generation, and prompt testing.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related