Qwen 3.8 2.4T
About this model
Qwen 3.8 2.4T, published as Qwen3.8-2.4T-A95B, is Alibaba's open-weight flagship language model, dated August 2026 in this catalog. It is a sparse mixture-of-experts system with roughly 2.4 trillion total parameters and about 95 billion active per token, a design that keeps inference cost tied to the active slice rather than the full weight count. NVIDIA's technical blog documents serving it on GB300 NVL72 systems, and quantized community checkpoints in NVFP4 format are published on Hugging Face, both signs that the model is intended for large multi-GPU rack deployment rather than single-node use.
Within the Qwen line, the scale jump is the headline change. The prior open flagship in this catalog, Qwen 3.5 397B, carried 397 billion total parameters, while Qwen 3.6 35B A3B targets lightweight deployment at 35 billion. Qwen 3.8 2.4T sits at the opposite end: far more total capacity, more active parameters per token, and a 262K-token context window, with the catalog description citing gains in software engineering, research and long-horizon agentic tasks over earlier releases.
Practical differences matter when choosing it. Input is text-only, so multimodal work belongs to siblings such as Qwen3 VL 235B. Thinking mode is mandatory in this deployment, meaning every response includes a reasoning pass — better for hard analysis and agentic chains, slower and more verbose for quick chat. The hosted Qwen 3.8 Max is the closed counterpart in the same generation.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 4d ago