minimaxiH3 Get access
xAIImage-to-videoVendor content policy applies

Grok Imagine Video v1.5 Developer Image-to-Video

xAI Grok Imagine Video v1.5 animates a starting frame image with natural-language motion prompts at 480p/720p/1080P.

$0.042per secondStarting price at the base resolution and quality tier.
$0.34Typical 8-second clip at the base tier. Longer or higher-resolution outputs scale with duration and tier.
Pay per useBilled per request from your balance. No subscription, no minimum.
Sample output · prompt: “Slow cinematic fly-through approaching a gigantic black hole. The camera begins with a wide shot of the surrounding galaxy, then gradually descends toward the glowing accretion disk. Massive rings of plasma rotate rapidl…”

What it does

Grok Imagine Video v1.5 Developer Image-to-Video, in practice.

Grok Imagine Video V1.5 is a frontier-tier image-to-video generation model developed by xAI that animates static images into short clips of up to 15 seconds with natively generated, synchronized audio — including dialogue, lip-sync, sound effects, and ambient music — produced in a single inference pass. This README applies to the following API model identifier: Released in preview around late May 2026, Grok Imagine Video V1.5 debuted at the top of the Artificial Analysis Video Arena Image-to-Video leaderboard with a 1404 ±6 Elo rating, surpassing ByteDance Seedance 2.0 and other established competitors. Built on xAI's Aurora engine — an autoregressive mixture-of-experts (MoE) network that jo

  • Native Synchronized Audio Generation: Audio (dialogue, lip-sync, SFX, ambient sound, music) is generated jointly with video tokens in a single inference pass rather than dubbed in post-processing. This produces event-aligned sound effects and natural lip-sync without requiring separate audio pipelines.
  • Aurora Autoregressive MoE Architecture: Unlike diffusion-transformer competitors, V1.5 uses an autoregressive mixture-of-experts network trained to predict next tokens from interleaved multimodal data. This unified token-space approach is what enables single-pass audio-video coherence.
  • Granular Duration Control (1–15 seconds): Clips can be requested at any integer second from 1 to 15, supporting precise targeting for short-form formats. V1.5 extends the prior 10-second limit by 50% while maintaining temporal coherence across the longer window.
  • Improved Physics and Photorealism: V1.5 introduces measurable gains in cloth dynamics, water simulation, hair motion, and object interaction. Subject deformation in high-motion scenes is reduced relative to V1.0, with sharper micro-expressions and improved translucent/glass material rendering.
  • Fast Inference: A 5-second 720p clip generates in approximately 20–30 seconds end-to-end — roughly 2–3× faster than Seedance 2.0.
  • Native 1080p Output: Clips render at 480p, 720p, or 1080p, with 1080p produced natively rather than upscaled from a lower-resolution pass.

Run Grok Imagine Video v1.5 Developer Image-to-Video

from $0.042/sec
Drop an image hereor click to choose a fileFiles stay in your project. Nothing is trained on.
resolution
aspect_ratio
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
promptNatural-language motion prompt. The starting frame is taken from the image.
image_urlPublic HTTPS URL or base64 data URI of the starting-frame image (JPEG, PNG, or WebP).
durationLength of generated video in seconds. Range: 1–15.default 8
resolutionOutput resolution.480p 720p 1080p
aspect_ratioOutput aspect ratio. The default matches the input image; specifying a different value stretches the image.1:1 16:9 9:16 4:3 3:4 3:2 2:3

Sample prompt

The prompt behind the sample.

Slow cinematic fly-through approaching a gigantic black hole. The camera begins with a wide shot of the surrounding galaxy, then gradually descends toward the glowing accretion disk. Massive rings of plasma rotate rapidly around the event horizon, while distant stars bend and warp through gravitational lensing. The camera subtly tilts and orbits around the black hole, emphasizing its immense scale. Tiny particles drift past the lens, creating depth and realism. Dynamic light scattering, cosmic dust trails, slow motion, breathtaking sci-fi spectacle, ultra realistic space environment.
duration: 8resolution: 720paspect_ratio: 9:16

FAQ

Short answers.

How much does Grok Imagine Video v1.5 Developer Image-to-Video cost?

Pricing starts at $0.042 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.34. Usage is billed per request from your balance — no subscription.

Does Grok Imagine Video v1.5 Developer Image-to-Video run uncensored here?

No. xAI applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Grok Imagine Video v1.5 Developer Image-to-Video for everything else it does well.

What does Grok Imagine Video v1.5 Developer Image-to-Video take as input?

It is a image-to-video model. xAI Grok Imagine Video v1.5 animates a starting frame image with natural-language motion prompts at 480p/720p/1080P.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run Grok Imagine Video v1.5 Developer Image-to-Video.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models