xAIxAI·🎬 Video Generation·New

Grok Imagine 1.5 Lite

private
Try on Venice.ai ↗
Quick reference
Grok Imagine 1.5 Lite — TLDR
  • 🆕 Image-to-video entry in xAI's Grok Imagine 1.5 generation.
  • 👁️ Animates a single still frame guided by a motion prompt.
  • 💬 Video and audio — effects, ambience, dialogue — generated in one pass.
  • ⚡ xAI reports the 1.5 Fast variant makes 6-second 720p clips in ~25s.
  • 📏 Native 1080p plus image and voice references added to the 1.5 line.
  • 🔧 Up to seven references per generation for consistent faces, voices.
  • 🔒 Proprietary xAI model; no open weights or published parameter count.
  • 🏢 Released in 2026 by xAI, a subsidiary of SpaceX.
💰 Pricing
$0.040 – $2.53
per generation
📅 On Venice since
Sep 30, 2026
2 days ago
Provider

xAI is an American artificial intelligence company and wholly owned subsidiary of SpaceX. The company develops AI systems under the Grok brand, spanning language models, image generation, and video synthesis. xAI has quickly established itself as a multimodal…

Read full profile →
26 models on Venice
9 video · 8 text · 4 image · 3 inpaint · 1 tts · 1 asr
Since Jan 29, 2026

About this model

Grok Imagine 1.5 Lite is the image-to-video entry in xAI's Imagine video line, released in 2026. You supply a starting frame plus a prompt describing motion, and the model animates the scene — camera moves, atmosphere and physics — while staying faithful to the source image. It sits alongside its text-driven counterpart Grok Imagine 1.5 Lite (text-to-video), and can be used alongside still-image generation from Grok Imagine 2.0.

Against the earlier generation represented by Grok Imagine, xAI reports improvements in audio generation, speech sync and motion stability: sound effects, ambience and dialogue are produced in the same pass and timed to the action, speech is clearer and better aligned, and motion holds together across a clip with fewer warps and more consistent weight and momentum. xAI also reports that its 1.5 Fast variant produces 6-second 720p videos in about 25 seconds, versus more than 40 seconds in the prior model; that figure is xAI's own number for the Fast variant and is not a measured result for this Lite endpoint.

The wider 1.5 family, including Grok Imagine 1.5 and Grok Imagine 1.5 R2V, subsequently gained native 1080p output plus image and voice references — up to seven per generation — so a character's face and voice can persist across shots.

In practice, prompts describe the motion and audio you want applied to the supplied frame, and longer sequences are assembled by chaining clips built from consistent reference frames.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 15h ago