Z.aiZ.ai·💬 Text Generation·New

GLM 5.3 Flash🔒Private

ReasoningVisionCodeFunction CallingWeb SearchE2EEprivate
🧠 Try in Intelligence →Try on Venice.ai ↗
Quick reference
GLM 5.3 Flash — TLDR
  • 🔒 Runs inside a Trusted Execution Environment with hardware attestation evidence
  • 🧠 320B-parameter MoE, roughly 18B active per token
  • 🆕 First natively multimodal model in Z.ai's GLM-5 series
  • 📏 1M-token context window, image and video input supported
  • ⚡ Hybrid sparse plus linear attention cuts attention compute and KV cache
  • 🔧 Function calling, web search, structured output, always-on reasoning
  • 📚 Trained on a 30-trillion-token multimodal corpus; MIT-licensed weights
  • 🎯 Aimed at coding agents, visual UI work, and long-horizon automation
💰 Pricing
$0.080 / $0.270
per 1M · input / output
📏 Context
1M tokens
📅 On Venice since
Sep 6, 2026
1 day ago
Provider

Z.ai, formally Knowledge Atlas Technology Joint Stock Co., Ltd., is a Chinese technology company specializing in artificial intelligence. Previously known internationally as Zhipu AI, the company rebranded to Z.ai in 2025. Its core focus is the GLM family of…

Read full profile →
15 models on Venice
14 text · 1 image
Since Apr 1, 2024

About this model

GLM 5.3 Flash is the efficiency tier of Z.ai's GLM-5 generation, offered here in an end-to-end encrypted deployment: the model runs inside a Trusted Execution Environment, and hardware attestation evidence can be independently verified to confirm enclave identity and configuration. Functionally it matches the standard GLM 5.3 Flash release, and it sits alongside the larger confidential-compute flagship GLM 5.3.

Architecturally, it is a mixture-of-experts model with about 320B total and 18B active parameters, routing each token through a small subset of experts in a stack that mixes linear-attention and sparse multi-head latent attention layers. Z.ai describes it as the first natively multimodal model in the GLM-5 series — vision is part of the 30-trillion-token pre-training corpus rather than a bolted-on encoder — accepting text, images and video and returning text.

Compared with same-family predecessors, the shift is both architectural and in training data. Z.ai reports that the hybrid sparse-plus-linear attention design reduces attention computation by 3.01 times and KV cache by 4.44 times relative to the dense-attention GLM 5.3, while preserving long-context quality, and that Manifold-Constrained Hyper-Connections further improve scaling efficiency. Z.ai reports higher benchmark scores than GLM 5.2 at lower serving cost.

Practical strengths follow from that design: a full 1M-token window for large repositories and long agent sessions, vision-driven UI coding from screenshots or screen recordings, tool and browser use, and structured JSON output. Weights for the underlying model are published under the MIT license.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Research & Papers

Primary reference paper for this model family, sourced from the HuggingFace model card.

Data sources: Venice API · HuggingFace · Wikipedia · arXiv — enrichment updated 1d ago