GoogleGoogle·💬 Text Generation·New

Gemini 3.7 Flash

ReasoningVisionFunction CallingWeb SearchAudioanonymized
🧠 Try in Intelligence →Try on Venice.ai ↗
Quick reference
Gemini 3.7 Flash — TLDR
  • - 🆕 Google's newest Flash-tier Gemini, released August 2026
  • - 📏 One-million-token context window for long documents and codebases
  • - 🧠 Tunable thinking budget to trade latency against answer depth
  • - 👁️ Multimodal input including text, images and audio
  • - 🔧 Function calling, tool use and multi-step agentic execution
  • - 🌐 Web-search grounding available for up-to-date answers
  • - 🏢 Hosted via Gemini API and Gemini Enterprise Agent Platform
  • - 🔒 Safety and capability evaluations published in a DeepMind model card
💰 Pricing
$1.88 / $9.38
per 1M · input / output
📏 Context
1M tokens
📅 On Venice since
Aug 14, 2026
3 days ago
Provider

Google is an American multinational technology corporation and one of the world's most valuable brands. A subsidiary of parent company Alphabet Inc., Google operates across search, cloud computing, consumer electronics, and artificial intelligence. Its…

Read full profile →
33 models on Venice
13 text · 11 video · 3 image · 3 inpaint · 1 music · 1 embedding · 1 tts
Since Oct 15, 2024

About this model

Gemini 3.7 Flash is Google's mid-tier "workhorse" entry in the Gemini 3 line, sitting between the deeper-reasoning Pro models and the lighter Flash-Lite tier. Google describes it as its most capable Flash model, built for complex coding, agentic workflows and reliable multi-step execution, with a one-million-token context window and tunable thinking. It accepts multimodal input, including text, images and audio, and supports function calling and web-search grounding.

It follows Gemini 3.6 Flash by roughly five weeks, continuing an unusually rapid cadence within the same family: Gemini 3 Flash Preview arrived in December 2025 and Gemini 3.5 Flash in May 2026. Across that sequence the headline context length has stayed at one million tokens, with the generational work concentrated on coding, tool use and sustained agentic behaviour rather than on raw context expansion. Google publishes a dedicated model card for 3.7 Flash through DeepMind, covering intended uses and safety evaluations.

Within Google's current lineup, Gemini 3.5 Flash-Lite remains the lower-latency option for simple high-volume tasks, while Gemini 3.1 Pro Preview targets harder reasoning problems. Gemini 3.7 Flash is offered as a hosted API model — through the Gemini API and the Gemini Enterprise Agent Platform — rather than as open weights, so parameter counts, quantization and licensing details are not disclosed. Developers can dial the thinking level per request to balance response speed against reasoning depth.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 3d ago