minimaxiH3 Get access
xAIText-to-videoVendor content policy applies

Grok Imagine Video v1.5 Text-to-Video

xAI Grok Imagine Video v1.5 generates video with native synchronized audio from a text prompt alone. Up to 15s at 480p, 720p, or 1080p.

$0.12per secondStarting price at the base resolution and quality tier.
$0.96Typical 8-second clip at the base tier. Longer or higher-resolution outputs scale with duration and tier.
Pay per useBilled per request from your balance. No subscription, no minimum.
Sample output · prompt: “Low angle wide shot of a massive, monolithic dark gray brutalist pyramid temple sitting in an infinite desert storm. Dust blowing across the frame. Massive, sleek matte-black spherical space vessels descending slowly fro…”

What it does

Grok Imagine Video v1.5 Text-to-Video, in practice.

Grok Imagine Video V1.5 is a frontier-tier video generation model developed by xAI that produces short clips of up to 15 seconds with natively generated, synchronized audio — including dialogue, lip-sync, sound effects, and ambient music — in a single inference pass. This README applies to the following API model identifier: Text-to-video is the prompt-only mode: no input image or reference material is required. The model composes subject, motion, camera work, and the full audio bed from the prompt alone, and renders natively at up to 1080p. Built on xAI's Aurora engine — an autoregressive mixture-of-experts (MoE) network that jointly models text, image, video, and audio tokens — the model r

  • Prompt-Only Generation: No source image or reference asset is needed. Scene, motion, camera behaviour, and audio are all derived from the text prompt, making this the fastest path from idea to finished clip.
  • Native Synchronized Audio Generation: Audio (dialogue, lip-sync, SFX, ambient sound, music) is generated jointly with video tokens in a single inference pass rather than dubbed in post-processing. This produces event-aligned sound effects and natural lip-sync without requiring separate audio pipelines.
  • Native 1080p Output: Text-to-video renders at 480p, 720p, or 1080p, with 1080p produced natively rather than upscaled from a lower-resolution pass.
  • Aurora Autoregressive MoE Architecture: Unlike diffusion-transformer competitors, V1.5 uses an autoregressive mixture-of-experts network trained to predict next tokens from interleaved multimodal data. This unified token-space approach is what enables single-pass audio-video coherence.
  • Granular Duration Control (1–15 seconds): Clips can be requested at any integer second from 1 to 15, supporting precise targeting for short-form formats.
  • Broad Format Support: Outputs H.264 MP4 at 24 FPS across seven aspect ratios (1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3).

Run Grok Imagine Video v1.5 Text-to-Video

from $0.12/sec
resolution
aspect_ratio
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
promptNatural-language description of the video to generate, including subject, motion, camera work, and audio direction. Native synchronized audio is generated in the same pass.
durationLength of generated video in seconds. Range: 1–15.default 8
resolutionOutput resolution.480p 720p 1080p
aspect_ratioOutput aspect ratio.1:1 16:9 9:16 4:3 3:4 3:2 2:3

Sample prompt

The prompt behind the sample.

Low angle wide shot of a massive, monolithic dark gray brutalist pyramid temple sitting in an infinite desert storm. Dust blowing across the frame. Massive, sleek matte-black spherical space vessels descending slowly from the foggy sky. Cinematic lighting, dramatic shadows, minimal aesthetic, high contrast, atmospheric fog, Denis Villeneuve style, hyper-realistic, slow camera tilt up, 4k.
duration: 5resolution: 720paspect_ratio: 16:9

FAQ

Short answers.

How much does Grok Imagine Video v1.5 Text-to-Video cost?

Pricing starts at $0.12 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.96. Usage is billed per request from your balance — no subscription.

Does Grok Imagine Video v1.5 Text-to-Video run uncensored here?

No. xAI applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Grok Imagine Video v1.5 Text-to-Video for everything else it does well.

What does Grok Imagine Video v1.5 Text-to-Video take as input?

It is a text-to-video model. xAI Grok Imagine Video v1.5 generates video with native synchronized audio from a text prompt alone. Up to 15s at 480p, 720p, or 1080p.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run Grok Imagine Video v1.5 Text-to-Video.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models