Qwen 3.8 2.4T

ReasoningCodeFunction CallingWeb Searchprivate
🧠 Try in Intelligence →Try on Venice.ai ↗
Quick reference
Qwen 3.8 2.4T — TLDR
  • 🏢 Alibaba open-weight flagship: 2.4 trillion total parameters, 95B active
  • 🧠 Sparse mixture-of-experts design; reasoning-first, thinking mode required here
  • 📏 262K-token context window; text-only input and output
  • 🎯 Aimed at software engineering, research and long-horizon agentic work
  • 🔧 Tooling support includes web search in this catalog's serving setup
  • ⚡ Served on NVIDIA GB300 NVL72 racks per NVIDIA's engineering blog
  • 🆕 Largest Qwen text model in this catalog by total parameters
  • 🔒 Released under a custom license rather than a standard open license
💰 Pricing
$2.50 / $7.50
per 1M · input / output
📏 Context
262K tokens
📅 On Venice since
Aug 12, 2026
5 days ago
Provider

Alibaba Group is a Chinese multinational technology company founded in 1999 and headquartered in Hangzhou, Zhejiang. Originally built around e-commerce and cloud computing, Alibaba has become one of the most prolific contributors to open-weight AI research,…

Read full profile →
61 models on Venice
23 video · 21 text · 7 image · 6 inpaint · 2 embedding · 2 tts
Since Jan 11, 2025

About this model

Qwen 3.8 2.4T, published as Qwen3.8-2.4T-A95B, is Alibaba's open-weight flagship language model, dated August 2026 in this catalog. It is a sparse mixture-of-experts system with roughly 2.4 trillion total parameters and about 95 billion active per token, a design that keeps inference cost tied to the active slice rather than the full weight count. NVIDIA's technical blog documents serving it on GB300 NVL72 systems, and quantized community checkpoints in NVFP4 format are published on Hugging Face, both signs that the model is intended for large multi-GPU rack deployment rather than single-node use.

Within the Qwen line, the scale jump is the headline change. The prior open flagship in this catalog, Qwen 3.5 397B, carried 397 billion total parameters, while Qwen 3.6 35B A3B targets lightweight deployment at 35 billion. Qwen 3.8 2.4T sits at the opposite end: far more total capacity, more active parameters per token, and a 262K-token context window, with the catalog description citing gains in software engineering, research and long-horizon agentic tasks over earlier releases.

Practical differences matter when choosing it. Input is text-only, so multimodal work belongs to siblings such as Qwen3 VL 235B. Thinking mode is mandatory in this deployment, meaning every response includes a reasoning pass — better for hard analysis and agentic chains, slower and more verbose for quick chat. The hosted Qwen 3.8 Max is the closed counterpart in the same generation.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 4d ago