Gemini 3.5 Flash-Lite
About this model
Gemini 3.5 Flash-Lite, released in July 2026 alongside Gemini 3.6 Flash, is Google's most cost-efficient and lowest-latency model in the Gemini 3.5 generation. It targets high-volume, cost-sensitive workloads such as agentic retrieval, classification, extraction, and document processing, while retaining a 1M-token context window and native multimodal input across text, images, and audio. Google's documentation positions it as a fit for less complex Gemini 3.5 Flash workloads that prioritize throughput.
Compared with its direct predecessor, the earlier Gemini 3.1 Flash-Lite, this release continues the Flash-Lite line as the fastest tier of the family, now built on the newer 3.5 generation architecture. As measured by the independent evaluator Artificial Analysis, output throughput is reported at roughly 350 tokens per second.
A defining feature is configurable thinking levels: it defaults to minimal thinking to optimize speed and cost for latency-sensitive tasks, but can scale up reasoning effort when quality matters. It also carries the 3.5-series tool suite, including function calling and built-in search.
Relative to the heavier Gemini 3.5 Flash and the newer Gemini 3.6 Flash, Flash-Lite trades peak reasoning depth for throughput and lower cost, making it the family's default choice for everyday questions, summarization, and lightweight coding.
This About section is AI-generated from public sources (Claude Opus 4.8), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 1d ago