xAIxAI·🎬 Video Generation·New

Grok Imagine 1.5 R2V

private
Try on Venice.ai ↗
Quick reference
Grok Imagine 1.5 R2V — TLDR
  • 👁️ Reference-to-video generation from xAI's Grok Imagine 1.5 family
  • 🎯 Reference images guide style and subject without fixing first frame
  • 🆕 Brings 1.5-generation video quality to the reference-conditioned mode
  • ⚡ xAI reports 1.5 Fast makes 6-second 720p clips in ~25 seconds
  • 💬 Native synchronized audio generated in the same pass
  • 🔧 xAI cites better motion, physics and audio versus prior Imagine video
  • 📚 Served through the Imagine API rather than as open weights
  • 🏢 Built by xAI, released July 2026
💰 Pricing
$0.100 – $2.43
per generation
📅 On Venice since
Jul 29, 2026
3 days ago
Provider

xAI is an American artificial intelligence company and wholly owned subsidiary of SpaceX. The company develops AI systems under the Grok brand, spanning language models, image generation, and video synthesis. xAI has quickly established itself as a multimodal…

Read full profile →
18 models on Venice
7 video · 5 text · 2 image · 2 inpaint · 1 tts · 1 asr
Since Jan 29, 2026

About this model

Grok Imagine 1.5 R2V is xAI's reference-to-video model, the 1.5-generation successor to Grok Imagine R2V from April 2026. Reference-to-video differs from image-to-video: instead of locking the supplied picture as the opening frame, reference images steer the style, subject and content of the generated clip, giving looser but more controllable guidance. The same Imagine surface also exposes text-conditioned and image-conditioned video modes, so the reference mode is one option among several rather than a separate product.

The upgrade path mirrors the rest of the 1.5 line. For the Video 1.5 release, xAI reported better motion, better physics and better audio at higher speed than its previous Imagine video model, with the Fast variant producing six-second 720p output in roughly 25 seconds versus 40-plus seconds before. Those generation-side gains are what the reference-conditioned variant inherits, alongside natively generated synchronized audio — effects, ambience and speech — produced in the same pass as the picture.

It sits beside the text-to-video and image-to-video builds of Grok Imagine 1.5 and Grok Imagine 1.5, plus still-image siblings such as Grok Imagine High Quality (SOTA) and the editing model Grok Imagine High Quality. Like the rest of the family, it is a hosted, closed-weight model reachable through xAI's Imagine API rather than a downloadable checkpoint, so no parameter count or license terms are published.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 14h ago