MoonshotMoonshot·💬 Text Generation·New

Kimi K3🔒Private

ReasoningVisionCodeFunction CallingWeb SearchE2EEprivate
🧠 Try in Intelligence →Try on Venice.ai ↗
Quick reference
Kimi K3 — TLDR
  • - 🧠 2.8T-parameter Mixture-of-Experts with 104B active parameters
  • - 📏 One-million-token context window for whole-repo and long-document work
  • - 👁️ Native multimodality: text, images and video in one model
  • - 🔧 Kimi Delta Attention, Attention Residuals and Stable LatentMoE architecture
  • - 🎯 Built for long-horizon agentic coding and knowledge work
  • - ⚡ Configurable reasoning effort; results reported at max effort
  • - 🔒 Served end-to-end encrypted, with function calling and web search
  • - 📚 Open weights released under Moonshot's own Kimi K3 License
💰 Pricing
$3.75 / $18.75
per 1M · input / output
📏 Context
1M tokens
📅 On Venice since
Sep 8, 2026
1 day ago
Provider

Moonshot is an AI research lab known for developing the Kimi family of large language models. The organization has gained recognition for building capable reasoning-oriented models, with the Kimi line representing its flagship series of text generation…

Read full profile →
8 models on Venice
8 text
Since Jan 27, 2026

About this model

Kimi K3 is positioned as Moonshot AI's flagship open-weight release — a 2.8-trillion-parameter Mixture-of-Experts system that Moonshot describes on its Hugging Face model card as the first open "3T-class" model. NVIDIA's hosted model documentation lists 104B activated parameters, a 160K vocabulary, and a 1,048,576-token input context. This catalog entry is the end-to-end-encrypted deployment of that model, alongside the standard Kimi K3 endpoint and the latency-tuned Kimi K3 Fast.

Architecturally it is a clear break from the K2 line represented by Kimi K2.6, Kimi K2.6 and Kimi K2.5. K3 is built on Kimi Delta Attention and Attention Residuals, with Stable LatentMoE routing and a 401M-parameter MoonViT-V2 vision encoder that makes image and video understanding native rather than a bolted-on adapter.

The stated design target is long-horizon work: navigating large repositories, iterating against logs, tests and screenshots, and producing research artifacts such as interactive dashboards and visualizations. Compared with the code-specialized Kimi K2.7 Code, K3 is a general frontier model rather than a coding-focused checkpoint.

Practical notes: the model is trained with preserved thinking history, so multi-turn and tool-calling applications must return prior assistant messages including reasoning content and tool calls. Reasoning effort is configurable, and Moonshot's published evaluations use the maximum setting.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 10h ago