xAI's image-to-video specialist: your still becomes the first frame — composition, light and identity preserved — with music, effects and lip-synced dialogue generated in the same pass. Type a prompt below and try it.
Developer
xAI
Type
Image-first · also T2V
Duration
Up to 15 s
Max output
720p (official)
Audio
Music · SFX · lip-sync, one pass
Speed
Renders in seconds
Arena
#1 i2v at launch (Elo 1473)
Our one-line verdict: the image-to-video pick. When the shot starts from a still you already have — a product photo, a portrait, a frame — nothing here animates it faster or more faithfully.
Grok Imagine Video 1.5 is the model that took the image-to-video crown in June: launched at the end of May, it debuted #1 on the crowd-sourced i2v arena at 1473 Elo — a 52-point jump over its predecessor, past Seedance 2.0 and Veo. Its architecture explains the specialty: Aurora generates frames sequentially, each conditioned on everything before it, so your uploaded still — treated as the literal first frame — keeps its composition, lighting and subject identity instead of drifting. Music, sound effects and lip-synced dialogue render in the same pass, and clips land in seconds rather than minutes. One disambiguation before you spend: this is xAI's video model, not the Grok chatbot — same brand, different product.
Judged from our own renders and the published record, revised as the facts move — that's a feature of this page, not a disclaimer.
• Image fidelity by architecture.
The source still is the first frame, not a loose reference — Aurora's sequential conditioning holds composition, light and identity through the whole clip. The i2v arena crown was won on exactly this.
• One-pass audio, dialogue included.
Background music, sound effects and lip-synced speech generate with the picture. Lip-sync is the headline improvement over 1.0 — a talking-head clip needs no second pass.
• Timestamped shot control.
Write the prompt as a timeline — '(0-5s) medium shot... (5-10s) push-in...' — and the beats land where you put them. Shot-list prompting, native.
• Speed as a workflow feature.
Renders typically finish in seconds. Iteration loops that cost minutes elsewhere run at conversation pace here — the cheapest quality multiplier there is.
• 720p official ceiling.
xAI's own line is 720p (some gateways expose higher). In a field where rivals ship 1080p to 4K, this is the spec that ends some briefs early.
• Text-to-video is the weaker half.
xAI's own guidance: t2v is 'more generative and less controlled' than the image path. If you don't have a starting still, the prompt-first flagships fit better.
• The brand-confusion tax.
It shares a name with the Grok chatbot but is a separate product — chatbot reasoning tells you nothing about this model, in either direction. Judge it on renders.
• Arena rank ≠ production record.
The Elo measures crowd preference on i2v pairs, and the model is weeks old. Treat the crown as a strong signal, not a settled verdict.
Copy into the generator above, swap the brackets, render.
Upload: your product photo
"(0-4s) the [bottle] stands in soft dawn light as dust motes drift. (4-9s) slow push-in while a warm key light fades up from the left. (9-12s) settle on a close-up as the label catches the light. Quiet ambient score, one soft chime at the end."
Timestamps turn the prompt into a shot list, and the still-as-first-frame architecture keeps the product exactly as photographed. The demo beside the generator is this prompt.
Upload: a portrait photo
"She looks up from the notebook and says, warmly: 'Took me years to learn this one.' Then a small smile, back to writing. Room tone, pencil scratch, no music."
Quoted dialogue engages the lip-sync — the 1.5 generation's proudest fix. Keep lines short and let the prompt name the before-and-after action so the delivery has somewhere to land.
Generate a first clip → Extend from its final frame → "Continue: the camera keeps drifting right past the window; outside, the rain has started. The music carries over without a seam."
Extension conditions on the exact last frame, so motion, light and audio continue rather than reset. Chain two or three extensions and a 15-second ceiling quietly becomes a 40-second sequence.
The image-to-video podium as the June arena ranked it — the new leader against the two it passed. Same stills and prompts across all three, judged by what we'd ship. All three selectable in the generator above.
preview May 30–31, #1 i2v arena debut, GA in the xAI API by mid-June as grok-imagine-video-1.5. 15-second clips, one-pass audio with improved lip-sync, extension and reference guidance. The version in the generator above.
the 10-second first release that put xAI on the video map; 1.5's arena jump was measured against it.
xAI's autoregressive image backbone, whose character-consistency strength is the architectural reason the video line treats your still as ground truth.
xAI's video generation model — formally Grok Imagine Video 1.5, released late May 2026 and GA in the xAI API by mid-June. It's image-first: upload a still, describe the motion, and it animates the scene with the source as the literal first frame, generating music, sound effects and lip-synced dialogue in the same pass. Clips run up to 15 seconds and render in seconds.