Z.aiZ.ai·💬 Text Generation·↑ Newer: GLM 5.2

GLM 4.6

ReasoningFunction CallingWeb Searchfp4private
🧠 Try in Intelligence →Try on Venice.ai ↗
Quick reference
GLM 4.6 — TLDR
  • 🆕 Z.ai's open-weight upgrade over GLM-4.5 across coding and reasoning
  • 📏 Context window expanded from 128K to 200K tokens
  • ⚡ Over 30% more efficient token use than GLM-4.5
  • 🧠 Stronger reasoning with tool use during inference
  • 🔧 Improved tool-using and search-based agentic performance
  • 🔒 Released openly under the permissive MIT License
  • 💬 Toggleable deep-thinking ("thinking") mode per request
💰 Pricing
$0.430 / $1.75
per 1M · input / output
📏 Context
198K tokens
📅 On Venice since
Apr 1, 2024
842 days ago
Provider

Z.ai, formally Knowledge Atlas Technology Joint Stock Co., Ltd., is a Chinese technology company specializing in artificial intelligence. Previously known internationally as Zhipu AI, the company rebranded to Z.ai in 2025. Its core focus is the GLM family of…

Read full profile →
12 models on Venice
11 text · 1 image
Since Apr 1, 2024

About this model

GLM 4.6 is a large language model from Z.ai (formerly Zhipu AI), whose GLM family of open-weight models is released under the MIT License. It positions itself as an agentic, reasoning, and coding foundation model, with the catalog listing reasoning, function-calling, and web-search capabilities and a context window near 200K tokens.

Compared to its same-family predecessor GLM-4.5, Z.ai reports several concrete improvements: the context window grew from 128K to 200K tokens for more complex agentic tasks, average token consumption dropped by over 30%, and reasoning now supports tool use during inference for stronger overall capability. The company also reports stronger performance in tool-using and search-based agents and better integration within agent frameworks versus GLM-4.5. Z.ai evaluated GLM-4.6 across eight public benchmarks and published its test questions and agent trajectories for reproduction.

GLM-4.6 supports a toggleable deep-thinking mode, enabled or disabled per request, building on the interleaved-thinking approach introduced with GLM-4.5.

Within the lineage, GLM-4.6 was succeeded by GLM 4.7. The family later expanded with GLM 5, GLM 5.1, and GLM 5.2, alongside lightweight variants like GLM 4.7 Flash.

This About section is AI-generated from public sources (Claude Opus 4.8), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Research & Papers

Primary reference paper for this model family, sourced from the HuggingFace model card.

Data sources: Venice API · HuggingFace · Wikipedia · arXiv — enrichment updated 5h ago