MiMo-V2.6-Flash
About this model
MiMo-V2.6-Flash is the efficiency-balanced entry in Xiaomi's MiMo-V2.6 generation, released in 2026. It is a sparse Mixture-of-Experts language model with roughly 309B total parameters of which about 15B are activated per token, so serving cost tracks the small active budget rather than the full weight count. The checkpoint natively understands text, images, video and audio in a single model, and accepts context windows up to one million tokens, which suits repository-scale code reading and long multi-session agent transcripts. Weights are published under an MIT license, and this deployment runs an FP8 compute path.
Relative to its same-family predecessor MiMo-V2.5, the V2.6 Flash release keeps Xiaomi's omnimodal direction while pushing the usable context to a million tokens and re-centering post-training on reinforcement learning. Per Xiaomi's model card, the V2.6 work is organised around scaling RL compute, environment diversity and grading compute together, using mixed reinforcement learning spanning coding, general agent, visual and cybersecurity tasks so that capability growth comes from exploration and feedback rather than pretraining scale alone.
In practice the model is positioned for long-horizon agentic work: iterative coding, tool use with function calling and web search, and computer-use style control loops, with reasoning available for harder multi-step problems. No independent third-party evaluation results for this checkpoint were available from trustworthy evaluators at the time of writing, so no benchmark figures are quoted here.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 3d ago