GoogleGoogle·💬 Text Generation·New

Gemini 3.8 Flash

ReasoningVisionFunction CallingWeb SearchAudioanonymized
🧠 Try in Intelligence →Try on Venice.ai ↗
Quick reference
Gemini 3.8 Flash — TLDR
  • - 🏢 Google's latest Gemini 3 Flash "workhorse" model, released September 2026.
  • - 📏 One million token context window; up to 64k output tokens.
  • - 🧠 Tunable thinking levels (low, medium, high) trade latency for accuracy.
  • - 👁️ Multimodal input: text, images, audio and video, plus code.
  • - 🔧 Built for agentic coding, tool use and long-horizon multi-step execution.
  • - 🎯 Google reports 3x more completed tasks than 3.7 Flash on document-heavy workflows.
  • - ⚡ Higher token consumption than 3.7 Flash; medium effort curbs it.
💰 Pricing
$0.938 / $4.69
per 1M · input / output
📏 Context
1M tokens
📅 On Venice since
Sep 2, 2026
1 day ago
Provider

Google is an American multinational technology corporation and one of the world's most valuable brands. A subsidiary of parent company Alphabet Inc., Google operates across search, cloud computing, consumer electronics, and artificial intelligence. Its…

Read full profile →
38 models on Venice
15 video · 14 text · 3 image · 3 inpaint · 1 music · 1 embedding · 1 tts
Since Oct 15, 2024

About this model

Gemini 3.8 Flash is the newest entry in Google's Flash line, positioned by Google as the "primary agentic workhorse" of the Gemini 3 family — sitting between the deep-reasoning Pro models and the high-throughput Flash-Lite tier such as Gemini 3.5 Flash-Lite. It handles text, images, audio and video with a one-million-token context window and up to 64k output tokens, and supports function calling and built-in tools for agent workflows.

Per its model card, Gemini 3.8 Flash is built directly on Gemini 3.7 Flash, which itself followed Gemini 3.6 Flash, Gemini 3.5 Flash and the original Gemini 3 Flash Preview. Google describes the generational gains as concentrated in software engineering, agentic tasks and multi-step reasoning in specialized domains, and says the model completed "more than three times as many tasks" as 3.7 Flash in its own evaluations of long-running, document-heavy workflows.

Google's developer guide notes the trade-off: relative to 3.7 Flash, 3.8 Flash delivers better accuracy and more reliable performance but consumes more tokens, which is why the effort-control mechanism matters. The same three thinking levels carried over from 3.7 Flash remain — low for latency-sensitive pipelines, medium as the default for complex code and agentic use, high for the hardest reasoning — with Google noting that medium effort still handles complex agentic tasks while trimming token use.

For deep-reasoning work, the Pro branch — represented here by Gemini 3.1 Pro Preview — remains the alternative track within the same generation.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 6h ago