Kimi K3🔒Private
About this model
Kimi K3 is positioned as Moonshot AI's flagship open-weight release — a 2.8-trillion-parameter Mixture-of-Experts system that Moonshot describes on its Hugging Face model card as the first open "3T-class" model. NVIDIA's hosted model documentation lists 104B activated parameters, a 160K vocabulary, and a 1,048,576-token input context. This catalog entry is the end-to-end-encrypted deployment of that model, alongside the standard Kimi K3 endpoint and the latency-tuned Kimi K3 Fast.
Architecturally it is a clear break from the K2 line represented by Kimi K2.6, Kimi K2.6 and Kimi K2.5. K3 is built on Kimi Delta Attention and Attention Residuals, with Stable LatentMoE routing and a 401M-parameter MoonViT-V2 vision encoder that makes image and video understanding native rather than a bolted-on adapter.
The stated design target is long-horizon work: navigating large repositories, iterating against logs, tests and screenshots, and producing research artifacts such as interactive dashboards and visualizations. Compared with the code-specialized Kimi K2.7 Code, K3 is a general frontier model rather than a coding-focused checkpoint.
Practical notes: the model is trained with preserved thinking history, so multi-turn and tool-calling applications must return prior assistant messages including reasoning content and tool calls. Reasoning effort is configurable, and Moonshot's published evaluations use the maximum setting.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 10h ago