GoogleGoogle·💬 Text Generation·New

Gemini 3.5 Flash-Lite

ReasoningVisionFunction CallingWeb SearchAudioanonymized
🧠 Try in Intelligence →Try on Venice.ai ↗
Quick reference
Gemini 3.5 Flash-Lite — TLDR
  • - ⚡ Fastest, lowest-cost model in Google's Gemini 3.5 family.
  • - 📏 Handles a 1M-token context window.
  • - 👁️ Multimodal input across text, images, and audio.
  • - 🧠 Adjustable thinking levels, defaulting to minimal for speed.
  • - 🔧 Native tool use including function calling and search.
  • - 🎯 Aimed at everyday questions, summarization, and lightweight coding.
  • - 🆕 Reported near 350 output tokens/s by Artificial Analysis.
  • - 🏢 Available for production on Google Cloud.
💰 Pricing
$0.375 / $3.13
per 1M · input / output
📏 Context
1M tokens
📅 On Venice since
Jul 9, 2026
13 days ago
Provider

Google is an American multinational technology corporation and one of the world's most valuable brands. A subsidiary of parent company Alphabet Inc., Google operates across search, cloud computing, consumer electronics, and artificial intelligence. Its…

Read full profile →
32 models on Venice
12 text · 11 video · 3 image · 3 inpaint · 1 music · 1 embedding · 1 tts
Since Oct 15, 2024

About this model

Gemini 3.5 Flash-Lite, released in July 2026 alongside Gemini 3.6 Flash, is Google's most cost-efficient and lowest-latency model in the Gemini 3.5 generation. It targets high-volume, cost-sensitive workloads such as agentic retrieval, classification, extraction, and document processing, while retaining a 1M-token context window and native multimodal input across text, images, and audio. Google's documentation positions it as a fit for less complex Gemini 3.5 Flash workloads that prioritize throughput.

Compared with its direct predecessor, the earlier Gemini 3.1 Flash-Lite, this release continues the Flash-Lite line as the fastest tier of the family, now built on the newer 3.5 generation architecture. As measured by the independent evaluator Artificial Analysis, output throughput is reported at roughly 350 tokens per second.

A defining feature is configurable thinking levels: it defaults to minimal thinking to optimize speed and cost for latency-sensitive tasks, but can scale up reasoning effort when quality matters. It also carries the 3.5-series tool suite, including function calling and built-in search.

Relative to the heavier Gemini 3.5 Flash and the newer Gemini 3.6 Flash, Flash-Lite trades peak reasoning depth for throughput and lower cost, making it the family's default choice for everyday questions, summarization, and lightweight coding.

This About section is AI-generated from public sources (Claude Opus 4.8), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 1d ago