About this model
Gemini Omni Flash 1.1 is the text-to-video entry point for Google's Omni line, a multimodal family that pairs Gemini's reasoning with DeepMind's generative media stack to produce video (with synchronized audio) alongside text in a single model. From a written prompt it renders a clip grounded in Gemini's real-world knowledge, with physics understanding intended to make motion and interactions more coherent.
Compared with the original Gemini Omni Flash, the 1.1 release is described by Google as a production-ready update centered on control. The most concrete change: when continuing a scene, the model now conditions on up to 10 seconds of prior context, where earlier models referenced only the final second, which Google credits with better visual consistency and narrative adherence. Videos can be extended in 10-second increments to a 40-second total, with the last 10 seconds used as context to keep motion, characters and audio coherent. Also new are explicit camera-movement control, first-and-last-frame interpolation, 4K upscaling, and a fast low-resolution draft mode for cheaper iteration.
The same 1.1 generation covers other input modes: image-to-video, R2V for reference-guided generation, and Edit for conversational revision of existing footage.
Google positions Omni as conversational video creation and editing, applies content safety filters to both prompts and generated video, and notes that English is fully supported while other languages remain unevaluated.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 1d ago