About this model
Grok Imagine 1.5 Lite is the image-to-video entry in xAI's Imagine video line, released in 2026. You supply a starting frame plus a prompt describing motion, and the model animates the scene — camera moves, atmosphere and physics — while staying faithful to the source image. It sits alongside its text-driven counterpart Grok Imagine 1.5 Lite (text-to-video), and can be used alongside still-image generation from Grok Imagine 2.0.
Against the earlier generation represented by Grok Imagine, xAI reports improvements in audio generation, speech sync and motion stability: sound effects, ambience and dialogue are produced in the same pass and timed to the action, speech is clearer and better aligned, and motion holds together across a clip with fewer warps and more consistent weight and momentum. xAI also reports that its 1.5 Fast variant produces 6-second 720p videos in about 25 seconds, versus more than 40 seconds in the prior model; that figure is xAI's own number for the Fast variant and is not a measured result for this Lite endpoint.
The wider 1.5 family, including Grok Imagine 1.5 and Grok Imagine 1.5 R2V, subsequently gained native 1080p output plus image and voice references — up to seven per generation — so a character's face and voice can persist across shots.
In practice, prompts describe the motion and audio you want applied to the supplied frame, and longer sequences are assembled by chaining clips built from consistent reference frames.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 15h ago