minimaxiH3 Get access

Models

Text-to-video models.

Type a shot, get a moving clip. Every text-to-video model here, with per-second pricing and a sample output. 71 models, priced per second, with a sample output each.

All 328Text-to-video 71Image-to-video 111Video-to-video 18Audio-to-video 3Text-to-image 60Image editing 65
MiniMax · Text-to-video

MiniMax H3 Max Text-to-Video

MiniMax H3 Max text-to-video: generate a cinematic video from a text prompt. Supports 480P、768P, 5-15s., and 16:9/9:16/1…

from $0.072/sec · spicy route
MiniMax · Text-to-video

MiniMax H3 Fast Text-to-Video

MiniMax H3 Fast text-to-video: generate a cinematic video from a text prompt. Supports 480P, 5-15s., and 16:9/9:16/1:1/a…

from $0.066/sec · spicy route
MiniMax · Text-to-video

MiniMax H3-Developer Text-to-Video

MiniMax H3-Developer self-hosted text-to-video: generate a video (with audio) from a text prompt. Supports 480P/768P/2K,…

from $0.030/sec · spicy route
MiniMax · Text-to-video

MiniMax H3 Text-to-Video

MiniMax H3 text-to-video: generate a cinematic video from a text prompt. Supports 2K, 5-15s., and 16:9/9:16/1:1/adaptive…

from $0.057/sec · spicy route
MiniMax · Text-to-video

MiniMax H3 Max Turbo Text-to-Video

MiniMax H3 Max Turbo text-to-video: generate a cinematic video from a text prompt. Supports 480P, 5-15s., and 16:9/9:16/…

from $0.036/sec · spicy route
Google · Text-to-video

Gemini Omni 1.1 Flash Reference-to-Video

A natively multimodal Google DeepMind model that generates cinematic, natively sound-enabled videos from a text prompt p…

from $0.055/sec
Google · Text-to-video

Gemini Omni 1.1 Flash Text-to-Video

A natively multimodal Google DeepMind model that turns a single text prompt into a cinematic clip with synchronized nati…

from $0.055/sec
Qwen · Text-to-video

Wan-3.0-Prime Text-to-video

All-in-one Wan3.0 renderer: cinematic, hyper-real video from a text prompt, up to 30s with smart-duration and adaptive a…

from $0.091/sec
Qwen · Text-to-video

Wan-3.0 Text-to-video

All-in-one Wan3.0 renderer: cinematic, hyper-real video from a text prompt, up to 30s with smart-duration and adaptive a…

from $0.060/sec
ByteDance · Text-to-video

Seedance 2.5 Text-to-Video

Generate videos from text prompts with native audio and optional web search.

from $0.20/sec
ByteDance · Text-to-video

Seedance 2.0 Mini Text-to-Video

Lightweight, economical video generation from text prompts with native audio.

from $0.017/sec
Qwen · Text-to-video

HappyHorse-1.1 Text-to-video

Generates videos from text prompts with HappyHorse 1.1, supporting 480P, 720P, or 1080P output, flexible aspect ratios, …

from $0.11/sec
Qwen · Text-to-video

HappyHorse-1.1 Reference-to-video

Generates videos from one to nine reference images and a text prompt, supporting 480P, 720P, or 1080P output, flexible a…

from $0.11/sec
Google · Text-to-video

Gemini Omni Flash Reference-to-Video

A natively multimodal Google DeepMind model that generates cinematic, sound-enabled videos from a text prompt plus 1-5 r…

from $0.20/sec
Google · Text-to-video

Gemini Omni Flash Text-to-Video

A natively multimodal Google DeepMind model that generates cinematic videos with synchronized native audio from a text p…

from $0.19/sec
Kuaishou · Text-to-video

Kling V3.0 Turbo Text-to-Video

Kling V3.0 Turbo Text-to-Video generates dynamic cinematic videos from text prompts using MVL technology. Supports first…

from $0.14/sec
Kuaishou · Text-to-video

Kling Video O3 4K Text-to-Video

Kling Omni Video O3 (4K) is Kuaishou advanced unified multi-modal video model with MVL (Multi-modal Visual Language) tec…

from $0.54/sec
xAI · Text-to-video

Grok Imagine Video v1.5 Developer Text-to-Video

xAI Grok Imagine Video v1.5 generates video with native synchronized audio from a text prompt alone. Up to 15s at 480p, …

from $0.042/sec
xAI · Text-to-video

Grok Imagine Video v1.5 Text-to-Video

xAI Grok Imagine Video v1.5 generates video with native synchronized audio from a text prompt alone. Up to 15s at 480p, …

from $0.12/sec
Google · Text-to-video

Gemini Omni Flash Text-to-Video Developer

Gemini Omni Flash is Google's multimodal video generation model. This text-to-video variant generates high-quality cinem…

from $0.17/sec
Qwen · Text-to-video

HappyHorse-1.0 Text-to-video

Generates videos from text prompts with HappyHorse 1.0, supporting 720P or 1080P output, flexible aspect ratios, and dur…

from $0.21/sec
Qwen · Text-to-video

HappyHorse-1.0 Reference-to-video

Generates videos from one to nine reference images and a text prompt, supporting 720P or 1080P output, flexible aspect r…

from $0.21/sec
ByteDance · Text-to-video

Seedance 2.0 Text-to-Video

Generate videos from text prompts with native audio and optional web search.

from $0.14/sec
ByteDance · Text-to-video

Seedance 2.0 Fast Text-to-Video

Fast video generation from text prompts with native audio.

from $0.041/sec
Qwen · Text-to-video

Wan-2.7 Text-to-video

Generates videos from text prompts with multi-shot narrative, audio generation, and sound-image synchronization.

from $0.15/sec
Google · Text-to-video

Veo 3.1 Lite Text-to-video

High-efficiency Veo 3.1 Lite text-to-video: create video with synchronized audio from text prompts. Targets high-volume …

from $0.075/sec
Google · Text-to-video

Veo3.1 Fast Text-to-video

Generate visually compelling videos from text in record time. Veo 3.1 Fast Text-to-Video prioritizes speed and responsiv…

from $0.12/sec
Google · Text-to-video

Veo3.1 Text-to-video

Generate high-fidelity videos from text prompts with Google’s most advanced generative video model. Veo 3.1 delivers cin…

from $0.30/sec
xAI · Text-to-video

Grok Imagine Video Text-to-Video

xAI Grok Imagine Video generates short videos (1-15s) from natural-language prompts at 480p or 720p.

from $0.075/sec
Vidu · Text-to-video

Vidu Q3-Turbo Text-to-video

Vidu Q3-Turbo Text-to-Video is an advanced AI video generation model that creates high-quality videos directly from text…

from $0.051/sec
Kuaishou · Text-to-video

Kling v3.0 Pro Text-to-Video

Kling v3.0 Professional Text-to-Video model by Kuaishou. Premium quality video generation from text prompts with advance…

from $0.14/sec
Kuaishou · Text-to-video

Kling v3.0 4K Text-to-Video

Kling v3.0 4K Text-to-Video model by Kuaishou. High-quality video generation from text prompts.

from $0.54/sec
Kuaishou · Text-to-video

Kling v3.0 Std Text-to-Video

Kling v3.0 Standard Text-to-Video model by Kuaishou. High-quality video generation from text prompts.

from $0.11/sec
Vidu · Text-to-video

Vidu Q3-Pro Text-to-video

Vidu Q3-Pro Text-to-Video is an advanced AI video generation model that creates high-quality videos directly from text d…

from $0.063/sec
ByteDance · Text-to-video

Seedance v1.5 Pro Text-to-Video

Native audio-visual joint generation model by ByteDance. Supports unified multimodal generation with precise audio-visua…

from $0.071/sec
Qwen · Text-to-video

Wan-2.6 Text-to-video

A speed-optimized text-to-video option that prioritizes lower latency while retaining strong visual fidelity. Ideal for …

from $0.11/sec
Kuaishou · Text-to-video

Kling Video O3 Pro Text-to-Video

Kling Omni Video O3 is Kuaishou's advanced unified multi-modal video model with MVL (Multi-modal Visual Language) techno…

from $0.14/sec
ByteDance · Text-to-video

Seedance v1.5 Pro Text-to-Video Fast

Native audio-visual joint generation model by ByteDance. Supports unified multimodal generation with precise audio-visua…

from $0.027/sec
Kuaishou · Text-to-video

Kling v2.6 Pro Text-to-Video

Latest text-to-video model from Kuaishou with sound generation, flexible aspect ratios, and cinematic quality.

from $0.090/sec
Kuaishou · Text-to-video

Kling Video O3 Std Text-to-Video

Kling Omni Video O3 (Standard) is Kuaishou's advanced unified multi-modal video model with MVL (Multi-modal Visual Langu…

from $0.11/sec
Kuaishou · Text-to-video

Kling Video O1 Text-to-video

Kling Omni Video O1 is Kuaishou's first unified multi-modal video model with MVL (Multi-modal Visual Language) technolog…

from $0.14/sec
PixVerse · Text-to-video

Pixverse c1 Reference-to-Video

Pixverse c1 Reference-to-Video model. High-quality video generation from image prompts.

from $0.045/sec
PixVerse · Text-to-video

Pixverse v6 Text-to-Video

Pixverse v6 Text-to-Video model. High-quality video generation from text prompts.

from $0.038/sec
PixVerse · Text-to-video

Pixverse v6 Reference-to-Video

Pixverse v6 Reference-to-Video model. High-quality video generation from image prompts.

from $0.038/sec
PixVerse · Text-to-video

Pixverse c1 Text-to-Video

Pixverse c1 Text-to-Video model. High-quality video generation from text prompts.

from $0.045/sec
Qwen · Text-to-video

Wan-2.5 Video Extend

Extend your videos with Alibaba WAN 2.5 video extender model with audio.

from $0.078/sec
MiniMax · Text-to-video

Hailuo-2.3 t2v Standard

High-quality text-to-video generation optimized for creative workflows with cinematic visuals and reliable prompt fideli…

from $0.42/sec
MiniMax · Text-to-video

Hailuo-2.3 t2v Pro

Professional-grade text-to-video model delivering advanced motion, physics realism and film-style output for VFX and mar…

from $0.73/sec
ByteDance · Text-to-video

Seedance v1 Pro Fast Text-to-video

An efficient text-to-video model geared toward fast, cost-effective generation. Ideal for prototyping short narrative cl…

from $0.013/sec
Kuaishou · Text-to-video

Kling v2.5 Turbo Pro Text-to-video

Delivers high-speed text-to-video generation with cinematic motion precision and enhanced temporal stability.

from $0.090/sec
Qwen · Text-to-video

Wan-2.5 Text-to-video Fast

Convert prompts into cinematic video clips with synchronized sound. Wan 2.5 generates 480p/720p/1080p outputs with stabl…

from $0.11/sec
Qwen · Text-to-video

Wan-2.5 Text-to-video

A speed-optimized text-to-video option that prioritizes lower latency while retaining strong visual fidelity. Ideal for …

from $0.053/sec
MiniMax · Text-to-video

Hailuo-02 t2v Pro

Hailuo 02 is a new AI video generation model from Hailuo AI.

from $0.73/sec
Kuaishou · Text-to-video

Kling v2.1 t2v Master

Interprets complex text prompts with advanced motion logic and enhanced dynamic-camera rendering.

from $0.36/sec
Kuaishou · Text-to-video

Kling v2.0 t2v Master

The foundational cinematic model combining high-fidelity visuals with realistic human motion generation.

from $0.36/sec
MiniMax · Text-to-video

Hailuo 02 t2v Standard

Hailuo 02 is a new AI video generation model from Hailuo AI.

from $0.42/sec
Kuaishou · Text-to-video

Kling v1.6 t2v Standard

Entry-level text-to-video generator offering stable motion and prompt alignment for short-form outputs.

from $0.072/sec
ByteDance · Text-to-video

Seedance v1 Pro t2v 1080p

A full-fidelity text-to-video model built for cinematic results. Generates multi-shot, 1080p videos with smooth motion, …

from $0.17/sec
ByteDance · Text-to-video

Seedance v1 Pro t2v 720p

A full-fidelity text-to-video model built for cinematic results. Generates multi-shot, 1080p videos with smooth motion, …

from $0.071/sec
ByteDance · Text-to-video

Seedance v1 Pro t2v 480p

A full-fidelity text-to-video model built for cinematic results. Generates multi-shot, 1080p videos with smooth motion, …

from $0.033/sec
Vidu · Text-to-video

Vidu Q2-Pro Reference-to-video

Vidu Q2-Pro Reference-to-Video is an advanced AI video generation model that brings static images to life. Upload a refe…

from $0.13/sec
Vidu · Text-to-video

Vidu Q2 Reference-to-video

Vidu Q2 Reference-to-Video is an advanced AI video generation model that brings static images to life. Upload a referenc…

from $0.096/sec
Vidu · Text-to-video

Vidu Q2 Text-to-video

Vidu Q2 Text-to-Video is an advanced AI video generation model that brings static images to life. Upload a reference ima…

from $0.063/sec
Vidu · Text-to-video

Vidu Q1 Reference-to-video

Vidu Q1 Reference-to-Video is an advanced AI video generation model that brings static images to life. Upload a referenc…

from $0.51/sec
Vidu · Text-to-video

Vidu Q1 Text-to-video

Vidu Q1 Text-to-Video is an advanced AI video generation model that brings static images to life. Upload a reference ima…

from $0.51/sec
Black Forest Labs · Text-to-video

BLACKFORESTLABS FLUX 3 Image-to-Video

Generate a video that starts from an input image using FLUX 3.

from $0.26/sec
Black Forest Labs · Text-to-video

BLACKFORESTLABS FLUX 3 Extend Video

Extend a video, continuing from the source clip final frames, using FLUX 3.

from $0.61/sec
Black Forest Labs · Text-to-video

BLACKFORESTLABS FLUX 3 Keyframes to Video

Generate a video that hits the supplied keyframe images at exact frame positions using FLUX 3.

from $0.26/sec
Black Forest Labs · Text-to-video

BLACKFORESTLABS FLUX 3 First & Last Frame to Video

Generate a video between a start frame and an end frame using FLUX 3.

from $0.26/sec
Black Forest Labs · Text-to-video

BLACKFORESTLABS FLUX 3 Text-to-Video

Generate a video (with audio) from a text prompt using FLUX 3.

from $0.26/sec
Lightricks · Text-to-video

Ltx 2.3 Quality Text-to-Video

Generate high-quality video with audio from images using LTX-2.3

from $0.0030/sec

Run any of them in Chat.

Invite code opens Chat with every model above. No code? Join the waitlist and tell us which model you need.

Get access Browse models