GoogleGoogle·🎬 Video Generation·New

Gemini Omni Flash 1.1 R2V

anonymized
Try on Venice.ai ↗
Quick reference
Gemini Omni Flash 1.1 R2V — TLDR
  • 🎬 Reference-to-video route of Google's Gemini Omni Flash 1.1 family
  • 🖼️ Combines text, images, audio and video references into one clip
  • 📏 Analyzes up to 10 seconds of prior context, versus one second before
  • 🎞️ Can reference up to three seconds of video for character consistency
  • 🔊 Generates synchronized audio — dialogue, ambience, effects — with the picture
  • 💬 Supports conversational, turn-by-turn editing that preserves unmentioned elements
  • 🧠 Grounded in Gemini world knowledge and physics reasoning for coherent motion
  • 🆕 Released August 2026 alongside text-, image- and video-to-video siblings
💰 Pricing
$0.160 – $3.50
per generation
📅 On Venice since
Aug 27, 2026
2 days ago
Provider

Google is an American multinational technology corporation and one of the world's most valuable brands. A subsidiary of parent company Alphabet Inc., Google operates across search, cloud computing, consumer electronics, and artificial intelligence. Its…

Read full profile →
37 models on Venice
15 video · 13 text · 3 image · 3 inpaint · 1 music · 1 embedding · 1 tts
Since Oct 15, 2024

About this model

Gemini Omni Flash 1.1 R2V is the reference-to-video entry point to Google's Gemini Omni Flash 1.1 generation, a fast multimodal family built for video generation and conversational video editing in the Gemini API. Rather than starting from a prompt alone, this route accepts combined multimodal references — text, images, audio and video together — to steer subject identity, motion, style and sound in the generated clip. Output includes synchronized audio produced with the picture, not dubbed afterward.

Compared with its direct predecessor Gemini Omni Flash R2V, the 1.1 update centers on control and continuity. Google states that Omni 1.1 can analyze up to 10 seconds of prior context, described as a leap from previous models that referenced only the final second, improving visual consistency and narrative adherence over longer stories. It also allows referencing up to three seconds of video when crafting a scene, so character appearance and visual context carry across shots.

The R2V route shares its base model with the other 1.1 variants — Gemini Omni Flash 1.1 for prompt-only generation, the image-to-video route, and Gemini Omni Flash 1.1 Edit — differing mainly in accepted inputs.

Like earlier Omni releases, it is positioned around Gemini's world knowledge, using physics, motion and cultural context to make interactions read coherently, and it supports iterative follow-up prompts where each turn builds on the previous video while preserving unmentioned elements.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 1d ago