About this model
Gemini Omni Flash 1.1 R2V is the reference-to-video entry point to Google's Gemini Omni Flash 1.1 generation, a fast multimodal family built for video generation and conversational video editing in the Gemini API. Rather than starting from a prompt alone, this route accepts combined multimodal references — text, images, audio and video together — to steer subject identity, motion, style and sound in the generated clip. Output includes synchronized audio produced with the picture, not dubbed afterward.
Compared with its direct predecessor Gemini Omni Flash R2V, the 1.1 update centers on control and continuity. Google states that Omni 1.1 can analyze up to 10 seconds of prior context, described as a leap from previous models that referenced only the final second, improving visual consistency and narrative adherence over longer stories. It also allows referencing up to three seconds of video when crafting a scene, so character appearance and visual context carry across shots.
The R2V route shares its base model with the other 1.1 variants — Gemini Omni Flash 1.1 for prompt-only generation, the image-to-video route, and Gemini Omni Flash 1.1 Edit — differing mainly in accepted inputs.
Like earlier Omni releases, it is positioned around Gemini's world knowledge, using physics, motion and cultural context to make interactions read coherently, and it supports iterative follow-up prompts where each turn builds on the previous video while preserving unmentioned elements.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 1d ago