About this model
Grok Imagine 1.5 R2V is xAI's reference-to-video model, the 1.5-generation successor to Grok Imagine R2V from April 2026. Reference-to-video differs from image-to-video: instead of locking the supplied picture as the opening frame, reference images steer the style, subject and content of the generated clip, giving looser but more controllable guidance. The same Imagine surface also exposes text-conditioned and image-conditioned video modes, so the reference mode is one option among several rather than a separate product.
The upgrade path mirrors the rest of the 1.5 line. For the Video 1.5 release, xAI reported better motion, better physics and better audio at higher speed than its previous Imagine video model, with the Fast variant producing six-second 720p output in roughly 25 seconds versus 40-plus seconds before. Those generation-side gains are what the reference-conditioned variant inherits, alongside natively generated synchronized audio — effects, ambience and speech — produced in the same pass as the picture.
It sits beside the text-to-video and image-to-video builds of Grok Imagine 1.5 and Grok Imagine 1.5, plus still-image siblings such as Grok Imagine High Quality (SOTA) and the editing model Grok Imagine High Quality. Like the rest of the family, it is a hosted, closed-weight model reachable through xAI's Imagine API rather than a downloadable checkpoint, so no parameter count or license terms are published.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 14h ago