Grok Imagine Video v1.5 Image-to-Video
xAI Grok Imagine Video v1.5 animates a starting frame image with natural-language motion prompts at 480p/720p/1080P.
What it does
Grok Imagine Video v1.5 Image-to-Video, in practice.
Grok Imagine Video V1.5 is a frontier-tier image-to-video generation model developed by xAI that animates static images into short clips of up to 15 seconds with natively generated, synchronized audio — including dialogue, lip-sync, sound effects, and ambient music — produced in a single inference pass. This README applies to the following API model identifier: Released in preview around late May 2026, Grok Imagine Video V1.5 debuted at the top of the Artificial Analysis Video Arena Image-to-Video leaderboard with a 1404 ±6 Elo rating, surpassing ByteDance Seedance 2.0 and other established competitors. Built on xAI's Aurora engine — an autoregressive mixture-of-experts (MoE) network that jo
- Native Synchronized Audio Generation: Audio (dialogue, lip-sync, SFX, ambient sound, music) is generated jointly with video tokens in a single inference pass rather than dubbed in post-processing. This produces event-aligned sound effects and natural lip-sync without requiring separate audio pipelines.
- Aurora Autoregressive MoE Architecture: Unlike diffusion-transformer competitors, V1.5 uses an autoregressive mixture-of-experts network trained to predict next tokens from interleaved multimodal data. This unified token-space approach is what enables single-pass audio-video coherence.
- Granular Duration Control (1–15 seconds): Clips can be requested at any integer second from 1 to 15, supporting precise targeting for short-form formats. V1.5 extends the prior 10-second limit by 50% while maintaining temporal coherence across the longer window.
- Improved Physics and Photorealism: V1.5 introduces measurable gains in cloth dynamics, water simulation, hair motion, and object interaction. Subject deformation in high-motion scenes is reduced relative to V1.0, with sharper micro-expressions and improved translucent/glass material rendering.
- Fast Inference: A 5-second 720p clip generates in approximately 20–30 seconds end-to-end — roughly 2–3× faster than Seedance 2.0.
- Native 1080p Output: Clips render at 480p, 720p, or 1080p, with 1080p produced natively rather than upscaled from a lower-resolution pass.
Run Grok Imagine Video v1.5 Image-to-Video
from $0.12/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | Natural-language motion prompt. The starting frame is taken from the image. | |
image_url | Public HTTPS URL or base64 data URI of the starting-frame image (JPEG, PNG, or WebP). | |
duration | Length of generated video in seconds. Range: 1–15. | default 8 |
resolution | Output resolution. | 480p 720p 1080p |
aspect_ratio | Output aspect ratio. The default matches the input image; specifying a different value stretches the image. | 1:1 16:9 9:16 4:3 3:4 3:2 2:3 |
Sample prompt
The prompt behind the sample.
Slow cinematic fly-through approaching a gigantic black hole. The camera begins with a wide shot of the surrounding galaxy, then gradually descends toward the glowing accretion disk. Massive rings of plasma rotate rapidly around the event horizon, while distant stars bend and warp through gravitational lensing. The camera subtly tilts and orbits around the black hole, emphasizing its immense scale. Tiny particles drift past the lens, creating depth and realism. Dynamic light scattering, cosmic dust trails, slow motion, breathtaking sci-fi spectacle, ultra realistic space environment.
duration: 8resolution: 720paspect_ratio: 9:16FAQ
Short answers.
How much does Grok Imagine Video v1.5 Image-to-Video cost?
Pricing starts at $0.12 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.96. Usage is billed per request from your balance — no subscription.
Does Grok Imagine Video v1.5 Image-to-Video run uncensored here?
No. xAI applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Grok Imagine Video v1.5 Image-to-Video for everything else it does well.
What does Grok Imagine Video v1.5 Image-to-Video take as input?
It is a image-to-video model. xAI Grok Imagine Video v1.5 animates a starting frame image with natural-language motion prompts at 480p/720p/1080P.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related