Gemini 3.7 Flash
About this model
Gemini 3.7 Flash is Google's mid-tier "workhorse" entry in the Gemini 3 line, sitting between the deeper-reasoning Pro models and the lighter Flash-Lite tier. Google describes it as its most capable Flash model, built for complex coding, agentic workflows and reliable multi-step execution, with a one-million-token context window and tunable thinking. It accepts multimodal input, including text, images and audio, and supports function calling and web-search grounding.
It follows Gemini 3.6 Flash by roughly five weeks, continuing an unusually rapid cadence within the same family: Gemini 3 Flash Preview arrived in December 2025 and Gemini 3.5 Flash in May 2026. Across that sequence the headline context length has stayed at one million tokens, with the generational work concentrated on coding, tool use and sustained agentic behaviour rather than on raw context expansion. Google publishes a dedicated model card for 3.7 Flash through DeepMind, covering intended uses and safety evaluations.
Within Google's current lineup, Gemini 3.5 Flash-Lite remains the lower-latency option for simple high-volume tasks, while Gemini 3.1 Pro Preview targets harder reasoning problems. Gemini 3.7 Flash is offered as a hosted API model — through the Gemini API and the Gemini Enterprise Agent Platform — rather than as open weights, so parameter counts, quantization and licensing details are not disclosed. Developers can dial the thinking level per request to balance response speed against reasoning depth.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 3d ago