MoonshotMoonshot·💬 Text Generation·New

Kimi K3 Fast

ReasoningVisionCodeFunction CallingWeb Searchprivate
🧠 Try in Intelligence →Try on Venice.ai ↗
Quick reference
Kimi K3 Fast — TLDR
  • 🏢 Moonshot AI's latest K-series generation, listed here as a separate endpoint.
  • 🧠 Open-weight multimodal reasoning model, per Moonshot's Hugging Face model card.
  • 📏 One-million-token context for whole repositories and large document sets.
  • 👁️ Accepts images alongside text.
  • 🆕 New attention stack: Kimi Delta Attention plus Attention Residuals, per the model card.
  • 🔧 Tool calling, web search and long-horizon agentic coding workflows.
  • 🎯 Aimed at repository navigation, debugging and iteration against logs and tests.
  • 🔒 Weights published publicly by Moonshot AI on Hugging Face.
💰 Pricing
$4.50 / $22.50
per 1M · input / output
📏 Context
1M tokens
📅 On Venice since
Aug 3, 2026
0 days ago
Provider

Moonshot is an AI research lab known for developing the Kimi family of large language models. The organization has gained recognition for building capable reasoning-oriented models, with the Kimi line representing its flagship series of text generation…

Read full profile →
5 models on Venice
5 text
Since Jan 27, 2026

About this model

Kimi K3 Fast is listed in this catalog as a separate speed-oriented endpoint of Moonshot AI's K3 model, the open-weight multimodal reasoning system the company publishes on Hugging Face; provider documentation specific to the Fast variant is not available, so its serving characteristics are not described here. The underlying K3 model, according to Moonshot's model card, accepts a one-million-token context window and takes images as well as text, and the catalog description positions it for complex coding, knowledge work and long-horizon agentic workflows rather than short chat turns.

The clearest generational change from the K2 line — including Kimi K2.5, Kimi K2.6 and the coding-focused Kimi K2.7 Code — is the attention stack. According to Moonshot's model card, K3 introduces Kimi Delta Attention, a hybrid linear attention mechanism, together with Attention Residuals, which allow representations to be retrieved selectively across depth. Both changes are directed at sustaining long contexts and extended tool-use loops. The standard Kimi K3 endpoint exposes the same released model family.

Typical uses described by the provider include navigating large codebases, cross-file refactors, tool- and terminal-driven agent runs, and research across long document sets, with screenshots, logs, test results and runtime feedback usable as inputs the model iterates against. Weights are released under Moonshot's own license terms rather than a standard open-source license.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 2h ago