XiaomiMiMoXiaomiMiMo·💬 Text Generation·New

MiMo-V2.6-Flash

ReasoningVisionCodeFunction CallingWeb SearchAudiofp8private
🧠 Try in Intelligence →Try on Venice.ai ↗
Quick reference
MiMo-V2.6-Flash — TLDR
  • 🧠 Sparse Mixture-of-Experts: 309B total parameters, 15B activated per token
  • 📏 One-million-token context for long repositories and extended agent traces
  • 👁️ Natively omnimodal: understands text, image, video and audio input
  • 🔧 Function calling, web search and computer-use oriented tool workflows
  • 🎯 Trained with mixed reinforcement learning across code, agents, vision, security
  • ⚡ Efficiency-balanced design; served in FP8 on this catalog
  • 🔒 Open weights released under the permissive MIT license
  • 🆕 Flash variant of the MiMo-V2.6 generation, released September 2026
💰 Pricing
$0.175 / $0.350
per 1M · input / output
📏 Context
1M tokens
📅 On Venice since
Sep 28, 2026
4 days ago
Provider

XiaomiMiMo is the large language model initiative from Xiaomi, the Chinese electronics and technology company, dedicated to developing capable open language models under the MiMo name. The effort reflects Xiaomi's broader push into foundational AI research…

Read full profile →
2 models on Venice
2 text
Since Jun 11, 2026

About this model

MiMo-V2.6-Flash is the efficiency-balanced entry in Xiaomi's MiMo-V2.6 generation, released in 2026. It is a sparse Mixture-of-Experts language model with roughly 309B total parameters of which about 15B are activated per token, so serving cost tracks the small active budget rather than the full weight count. The checkpoint natively understands text, images, video and audio in a single model, and accepts context windows up to one million tokens, which suits repository-scale code reading and long multi-session agent transcripts. Weights are published under an MIT license, and this deployment runs an FP8 compute path.

Relative to its same-family predecessor MiMo-V2.5, the V2.6 Flash release keeps Xiaomi's omnimodal direction while pushing the usable context to a million tokens and re-centering post-training on reinforcement learning. Per Xiaomi's model card, the V2.6 work is organised around scaling RL compute, environment diversity and grading compute together, using mixed reinforcement learning spanning coding, general agent, visual and cybersecurity tasks so that capability growth comes from exploration and feedback rather than pretraining scale alone.

In practice the model is positioned for long-horizon agentic work: iterative coding, tool use with function calling and web search, and computer-use style control loops, with reasoning available for harder multi-step problems. No independent third-party evaluation results for this checkpoint were available from trustworthy evaluators at the time of writing, so no benchmark figures are quoted here.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 3d ago