AlibabaAlibaba·💬 Text Generation·New

Qwen 3.8 Flash

ReasoningVisionCodeFunction CallingWeb Searchanonymized
🧠 Try in Intelligence →Try on Venice.ai ↗
Quick reference
Qwen 3.8 Flash — TLDR
  • 📏 Native 1M-token context window for long documents and codebases
  • 🧠 Thinking mode can be switched on or off per request
  • 👁️ Multimodal input: text, images and video
  • 🔧 Function calling and agentic tool-use workflows supported
  • ⚡ Speed-oriented tier of the Qwen 3.8 generation
  • 🌐 Open-weight Flash-Next checkpoint published on Hugging Face
  • 🎯 Positioned for coding assistance and visual understanding
  • 🏢 Released by Alibaba's Qwen team, September 2026
💰 Pricing
$0.140 / $0.490
per 1M · input / output
📏 Context
1M tokens
📅 On Venice since
Sep 10, 2026
10 days ago
Provider

Alibaba Group is a Chinese multinational technology company founded in 1999 and headquartered in Hangzhou, Zhejiang. Originally built around e-commerce and cloud computing, Alibaba has become one of the most prolific contributors to open-weight AI research,…

Read full profile →
75 models on Venice
35 video · 23 text · 7 image · 6 inpaint · 2 embedding · 2 tts
Since Jan 11, 2025

About this model

Qwen 3.8 Flash is the speed-oriented tier of Alibaba's Qwen 3.8 generation, sitting alongside the flagship Qwen 3.8 Max and the open-weight Qwen 3.8 2.4T, and above the compact Qwen 3.8 27B. Alibaba documents it as a multimodal model that pairs reasoning and generation with a natively supported million-token context window, so entire codebases, long document sets or extended conversations can be handled in a single pass.

Compared with earlier hosted tiers in the family such as Qwen 3.7 Plus, the Flash tier of this generation is built around that very long context together with multimodal input, accepting text, images and video rather than text alone. Reasoning behaviour is also more controllable: thinking mode is exposed as a per-request switch, so the same endpoint can serve low-latency chat and deliberate, step-by-step problem solving.

The model targets coding assistance, agentic workflows and visual understanding, with function calling available for tool-driven pipelines. An open-weight sibling checkpoint, Qwen3.8-Flash-Next, is published by the Qwen team on Hugging Face for teams that prefer to self-host rather than call the hosted endpoint.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 3d ago