Qwen 3.6 35B A3B

ReasoningVisionCodeFunction CallingWeb Searchfp8private
🧠 Try in Intelligence →Try on Venice.ai ↗
Quick reference
Qwen 3.6 35B A3B — TLDR
  • - 🧠 Sparse mixture-of-experts: 35B total, ~3B active per token.
  • - 📏 Native 256K-token context window.
  • - 🎯 Tuned for agentic coding, STEM reasoning, and tool use.
  • - 🔧 Built-in function calling and web-search capabilities.
  • - 🔒 Apache-2.0 licensed, self-hostable open weights.
  • - ⚡ Ships as an FP8 checkpoint for efficient inference.
  • - 🆕 Flagship open-weight release of the Qwen3.6 generation.
  • - 🏢 Developed by Alibaba's Qwen team.
💰 Pricing
$0.150 / $1.00
per 1M · input / output
📏 Context
256K tokens
📅 On Venice since
Jul 20, 2026
2 days ago
Provider

Alibaba Group is a Chinese multinational technology company founded in 1999 and headquartered in Hangzhou, Zhejiang. Originally built around e-commerce and cloud computing, Alibaba has become one of the most prolific contributors to open-weight AI research,…

Read full profile →
52 models on Venice
20 video · 19 text · 5 image · 4 inpaint · 2 embedding · 2 tts
Since Jan 11, 2025

About this model

Qwen 3.6 35B A3B is a text model from Alibaba's Qwen team and the flagship open-weight release of the Qwen3.6 generation. It uses a sparse mixture-of-experts design with 35 billion total parameters but only about 3 billion active per token, and natively supports a 256K-token context window. The weights ship under Apache-2.0, including an official FP8 checkpoint, and the model is tuned for agentic coding, STEM reasoning, and tool use.

Relative to its direct predecessor Qwen 3.5 35B A3B, the model keeps the same 35B-total / 3B-active mixture-of-experts layout, positioning the 3.6 release as a successor within the same size class rather than a parameter scale-up. It also sits alongside the smaller Qwen 3.6 27B within the same family.

The model targets developers building coding agents and tool-using workflows, with function-calling and web-search capabilities and open weights that can be self-hosted. As a compact active-parameter MoE, it aims to combine the throughput of a small model with the capacity of a larger expert pool, making it suitable for latency-sensitive agentic and STEM applications.

This About section is AI-generated from public sources (Claude Opus 4.8), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 23h ago