Our best video model, now with text, image, and voice references — generating up to 1080p.
When we launched Imagine Video 1.5 last month, it was our best video model yet — better motion, better physics, and better audio. Today it goes further: image and voice references, video from a prompt alone, and native 1080p generation.
Image and voice references start today in the US for SuperGrok Heavy and SuperGrok Plus on grok.com/imagine and iOS, rolling out to all tiers over the next few days.
Describe the shot — no starting image needed. Text-to-video pairs our image generation with image-to-video. Native 1080p is now supported with text-to-video and image-to-video. Text-to-video and native 1080p are generally available on grok.com/imagine, iOS, and Android.
Pass in a character image and a voice reference, and both hold — the same face and the same voice in every scene.

Character
Voice
Each reference image locks one thing in place — a face, a product, a location. Keep a character and swap the scene, keep the scene and swap the character, or hold both and change only the action. Up to seven references per generation.

Character

Scene

Character

Scene

Character

Scene
Image references, text-to-video, and native 1080p are live in the xAI API with our best video model, grok-imagine-video-1.5. Voice reference support is available on request.