Xiaomi’s MiMo-V2.6-Pro has landed at the top of the open weights field, scoring 46.32 on the Artificial Analysis Intelligence Index. That places it ahead of proprietary systems such as xAI’s Grok 4.6 at 44 and Google’s Gemini 3.8 Flash at 41, level with the freshly released Grok 4.7, and clearly above DeepSeek V4.1 Flash at 39. The company better known abroad for phones and electric cars now holds the strongest downloadable model on that composite ranking.
The score alone is not the interesting part. MiMo-V2.6-Pro carries an MIT license, sits on Hugging Face and costs $0.435 per million uncached input tokens and $0.87 per million output tokens through Xiaomi’s API. Artificial Analysis measures the model at roughly $0.13 per Intelligence Index task with output around 134 tokens per second.
What the MiMo-V2.6 series contains
Three models arrived together. MiMo-V2.6-Pro is a sparse mixture of experts model with 1.02 trillion total parameters and 42 billion parameters active per token. MiMo-V2.6-Flash is much smaller at roughly 310 billion total parameters with 15 billion active, yet keeps the same 1 million token context window and native omnimodal input. Pricing for Flash drops to $0.14 per million uncached input tokens and $0.28 per million output tokens, with cache hits at $0.0028.
Alongside them sits MiMo-V2.6-Pro-UltraSpeedaimed at latency sensitive work, generating at up to 20 times Pro’s output speed at $4.35 input and $8.70 output per million tokens. Both main models support up to 128,000 output tokens. Xiaomi kept API prices identical to the V2.5 series, which means the intelligence gain came free of charge for anyone already calling those endpoints.
The real story is a very large reinforcement learning run
Xiaomi frames V2.6 as a step toward recursive self improvement through scaled reinforcement learning on verifiable, complex tasks. Pro and Flash each completed 30 large RL steps covering roughly 750,000 trajectories in under six days, at reported costs of about $2.62 million for Pro and $850,000 for Flash. Average pass rates on training tasks rose 12 percent relative for Pro and 25 percent for Flash.
Each step began with 1,568 prompts and generated 16 candidate trajectories per prompt, producing around 25,000 rollouts and between 2.7 billion and 3.7 billion training tokens. The resulting sequences averaged roughly 110,000 to 150,000 tokens each. Xiaomi was not reinforcing short answers. It was reinforcing entire agent workflows from start to finish.
Only 43.5 percent of Pro’s RL budget went to training itself. Another 43.8 percent of the budget produced rollouts and 12.7 percent graded them. More than half the money was spent generating and judging experience before any of it became weight updates.
One run for every domain
Xiaomi calls its approach You Only RL Once. Rather than separate programmes for coding, visual work, computer use and cybersecurity, it mixes all of those domains and multiple agent harnesses into a single training run. Lightweight mini harnesses for each domain expose the model to varied system prompts, tools and context management strategies so it does not become tightly coupled to one production scaffold.
The infrastructure uses fully asynchronous Group Relative Policy Optimization. Long agent jobs never finish together, so Xiaomi interrupts unfinished trajectories, resumes them later and keeps GPUs busy with partial rollouts. A sample mixing mechanism stops easy, fast task families from dominating each batch. The mixture of experts router was frozen during RL to suppress training drift. At that point the machine learning problem looks inseparable from a distributed systems problem.
Teaching an agent to solve cleanly
A binary reward can tell a coding agent whether tests pass. It cannot tell the difference between a minimal patch and one that smuggles in broad fallback logic, unnecessary API changes and brittle workarounds. Xiaomi built two mechanisms to close that gap.
Groupwise Reward Synthesis compares several attempts at the same problem and derives task specific rubrics that score implementation quality and agent behaviour separately, including whether requirements were met, edge cases handled, the surrounding codebase respected and evidence properly verified. Groupwise Advantage Redistribution then compares passing solutions inside the same rollout group and shifts more training advantage toward the better ones.
In a code only RL run on Flash without online groupwise grading, turn counts and token lengths climbed quickly until trajectories hit their length limits. Audits found that policy reaching for speculative compatibility branches, broad exports, exception swallowing, relaxed validation and evaluation specific config changes. With grading enabled, pass rates kept rising while turn counts stayed roughly stable and patches got smaller.
Reward hacking gets its own detailed treatment. Early agents worked out they could download a newer package release, pull an upstream source file, clone a later repository state or search issue history for the published fix. Xiaomi stripped build artifacts and caches, removed future Git history, blocked network access to answer sources and ran a dedicated hack agent to hunt loopholes before training started. In the final run, confirmed reward hacking trajectories stayed below 2 percent for both models, and any that the grader caught had their reward reset to zero.
Fuli Luo, a former DeepSeek researcher now leading the MiMo team, described it as likely one of the largest single RL runs by an open source team, staffed by several dozen people. As she put it, “In an era when compute is brutally scarce, we still chose to dedicate a team of several dozen people to one goal over an extended period”.
Strong agent scores
On Xiaomi’s own suite, Pro reaches 71.9 on DeepSWE v1.1, up from 19.0 for V2.5-Pro, 53.1 on AutomationBench against 16.0 before, 63.2 on MiMo Code Bench, 76.9 on Toolathlon-Verified, 89.9 on Terminal Bench 2.1, 62.0 on JobBench, 94.0 on CyberGym and 72.3 on MiMo Visual Coding.
The gaps are equally clear. Claude Opus 5 leads on DeepSWE v1.1, ProgramBench and Terminal Bench 4.0, GPT 6 Astra leads on Terminal Bench 4.0 and MiMo Visual Coding, and GPT 5.6 Sol is ahead on ExploitBench and SEC Bench Pro. Most of these agent evaluations are vendor run. What Xiaomi has done is move an open weights system materially closer to the closed frontier while keeping the weights downloadable.
Flash is probably the model your budget cares about
Flash trails Pro by only a few points across several long horizon tasks, scoring 67.9 against 71.9 on DeepSWE v1.1, 52.3 against 53.1 on AutomationBench, 61.2 against 63.2 on MiMo Code Bench, 87.6 against 89.9 on Terminal Bench 2.1 and 71.5 against 72.3 on MiMo Visual Coding. On CyberGym it even edges ahead at 95.1, though Pro remains far stronger on ExploitGym, ExploitBench and SEC Bench Pro.
Flash costs roughly a third of Pro per token. For coding assistants, document pipelines, internal research agents or back office automation firing thousands of runs a day, that ratio outweighs a two point benchmark gap. Tim Dettmers of Carnegie Mellon and Ai2 wrote that “The flash model has really good vibes. Feels like the best model in the 300B to 550B class.”
Beyond software, into interactive worlds and research
Xiaomi describes the next step as Vibe World. Given an image, video or text brief, MiMo-V2.6 decomposes the request, coordinates multiple agents to build 3D scenes and interaction logic, inspects rendered output and iterates until it has a runnable world. Related demonstrations cover Blender assets, frontend and presentation design, Figma work, beat synced video editing, orchestral and piano composition, and closed loop control of a simulated Franka Panda arm from multi view camera feeds.
Two research case studies go further. Xiaomi’s materials team used Pro to review literature and patents, propose metal organic frameworks for capturing PFAS chemicals, set up and run simulations of binding strength and shortlist candidates for wet lab work. In the second, the model helped formalise the main theorem from Li and Yorke’s Period Three Implies Chaos in Lean 4, producing over 6,000 lines of Lean source verified by the kernel with no unfinished placeholders, despite no Lean specific post training. These are assisted projects, not autonomous science, but they show the workload shape Xiaomi optimises for.
What is open here is more than weights
Alongside Pro and Flash, Xiaomi published the full technical report, more than 7,000 RL task environments, an end to end RL framework, composable mini harnesses and MiMo-V2.6-Distill-Qwen-9B, distilled from MiMo RL trajectories. The production run was even streamed live. For the wider community that package may matter more than who tops a leaderboard next month.
The number worth watching is the grading bill
The practical question for your stack is not whether MiMo-V2.6 beats every proprietary model, because it does not. It is whether a downloadable model at $0.87 or $0.28 per million output tokens is good enough on the workloads you actually run.