Xiaomi's MiMo-V2.6-Pro Is the World's Best Open-Weight Model — and Anthropic Says Claude Helped Build It
Released under an MIT license just eleven days after Anthropic accused Xiaomi of illicitly distilling Claude's capabilities, MiMo-V2.6-Pro has debuted as the highest-scoring open-weight model on the Artificial Analysis Intelligence Index — and it's reshaping what developers expect from Chinese AI labs.
Sophia Chen🇨🇦 China Desk CorrespondentOct 8, 2026 10m readThe Smartphone Giant That Quietly Built a Frontier AI Lab
There is a version of the Chinese AI story that centres entirely on DeepSeek, Alibaba, and Zhipu — the labs that have dominated headlines since early 2025. Xiaomi does not fit neatly into that narrative. The company is best known for affordable smartphones, smart home devices, and electric vehicles. Its AI ambitions have been treated, until recently, as a sideshow to its hardware business.
That framing became harder to sustain on September 22, 2026, when Xiaomi released MiMo-V2.6-Pro — a 1.02-trillion-parameter sparse Mixture-of-Experts model that immediately claimed the top spot among open-weight systems on the Artificial Analysis Intelligence Index↗, scoring 46 points and tying with xAI's closed-source Grok 4.7. The weights are available on Hugging Face↗ under an MIT license, ungated and free to download. The API costs $0.435 per million input tokens and $0.87 per million output tokens — a fraction of what comparable proprietary models charge.
The release arrived eleven days after Anthropic published a threat intelligence report accusing Xiaomi of illicitly routing hundreds of thousands of user conversations through Claude to harvest training data. The timing was not lost on anyone watching.
From 7B to a Trillion Parameters in Eighteen Months
To understand what MiMo-V2.6-Pro represents, it helps to trace how quickly Xiaomi AI Lab has scaled. The company launched its first public language model — MiMo-7B, a compact reasoning model tuned for mathematics and coding — on April 30, 2025. The project was led by a small team and attracted little attention outside specialist circles.
Everything changed in late 2025 when Xiaomi recruited Luo Fuli, a former core developer at DeepSeek. Her arrival triggered a rapid scaling phase: by December 2025, the lab had shipped a 309B MoE architecture; by March 2026, a trillion-parameter flagship. Xiaomi has pledged approximately US$8.7 billion in AI R&D over three years, a commitment that has funded both the compute infrastructure and the talent pipeline needed to sustain this pace.
The MiMo series is not a standalone product. It is designed as the "neural connective tissue" for Xiaomi's Human × Car × Home ecosystem — embedded in HyperOS, routing simple queries to on-device quantized models and escalating complex reasoning tasks to cloud-based inference. That vertical integration gives Xiaomi a distribution advantage that pure-play AI labs cannot easily replicate.
"The emergence of MiMo-V2.6 — coinciding with Alibaba's V900 chip announcement — signals the maturing of an independent Chinese AI ecosystem that no longer depends on any single lab or any single piece of Western hardware." — Forkast News↗
What MiMo-V2.6-Pro Actually Is
The official model page↗ describes MiMo-V2.6-Pro as an omnimodal sparse MoE system. The key specifications:
- Total parameters: 1.02 trillion, with 42 billion active per inference step — a ratio that keeps compute costs manageable despite the headline scale
- Architecture: 70-layer design with a hybrid attention mechanism combining 128-token sliding-window layers and global attention layers, enabling efficient long-context processing
- Context window: 1 million tokens, designed for long-horizon agentic tasks
- Modalities: Native understanding of text, image, video, and audio inputs; text output
- Speculative decoding: A 5-layer multi-token prediction (MTP) decoder for faster parallel verification
- Variants: MiMo-V2.6-Flash (310B total / 15B active, optimised for high-volume production) and MiMo-V2.6-Pro-UltraSpeed (API-only, up to 20× faster output)
The model is deployable via vLLM and SGLang↗ with tensor parallelism. Self-hosting requires approximately 566 GB of disk space using native mxfp4 quantization, and typically demands an 8-GPU server-grade node. For most developers, the hosted API or OpenRouter will be the practical entry point.
The "You Only RL Once" Training Strategy
The training methodology is one of the more technically interesting aspects of the release. Rather than running separate reinforcement learning passes for each capability domain — a common approach that risks catastrophic forgetting — Xiaomi's team executed a single, massive mixed RL run across coding, visual tasks, cybersecurity, and general agent workflows simultaneously.
This "You Only RL Once" strategy used fully asynchronous Group Relative Policy Optimization (GRPO) and a novel Groupwise Reward Synthesis (GRS) mechanism to reduce reward hacking. The entire RL run consumed roughly 750,000 trajectories over six days at a reported cost of approximately $2.62 million for the Pro variant — a figure Xiaomi publicised via a live training dashboard, an unusual transparency move that drew both admiration and scepticism from the research community.
A subsequent Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD) stage, documented in the MOPD checkpoint on Hugging Face↗, addressed capability gaps in domains where automated reward design is unreliable — long-horizon game development, scientific research, embodied intelligence. The MOPD process also resolved a persistent agentic failure mode: tool-call repetition, where models issue identical tool calls in loops without making progress.
Benchmark Performance: Where It Leads, Where It Trails
On the BenchLM Chinese AI model leaderboard↗ as of October 7, 2026, MiMo-V2.6-Pro holds the top position with a composite score of 74.2, ahead of Moonshot AI's Kimi K3 (70.6) and Alibaba's Qwen3.8 Max (70.5). On the Artificial Analysis Intelligence Index, it scores 46, tying with Grok 4.7 and sitting below Claude Opus 5 (51) and GPT-6 Astra (53).
The model's strongest domains are agentic and automation tasks:
- DeepSWE v1.1 (software engineering): 71.9% — up from 19.0% on its predecessor MiMo-V2.5-Pro, a dramatic improvement
- Terminal Bench 2.1: 89.9%
- CyberGym (cybersecurity): 94.0%
- AutomationBench v1.0.6: 53.1% — outperforming Claude Opus 5 (50.3%) and GPT-5.6 Sol (45.8%)
- Toolathlon-verified (tool use): 76.9%
The picture is less flattering on harder, longer-horizon benchmarks. MiMo-V2.6-Pro scores 34.9 on Terminal-Bench 4.0, compared to Claude Opus 5's 49.0 and GPT-5.6's 39.9. On ExploitBench, it reaches 47.9 against Claude's 70.0 and GPT-5.6's 78.5. These gaps suggest the model excels at structured, verifiable agentic tasks but struggles with the kind of open-ended, multi-step reasoning that the most advanced proprietary systems handle more reliably.
"MiMo-V2.6-Pro is the first open-weight model to genuinely threaten the proprietary frontier on automation benchmarks — but the gap on hard, long-horizon tasks is real and should not be papered over." — The Decoder↗
Competitive Positioning Against Chinese Peers
Within the Chinese open-weight landscape, MiMo-V2.6-Pro's arrival reshuffles a competitive order that had been relatively stable since DeepSeek's V4 family completed its rollout in September 2026. The key comparisons:
- vs. DeepSeek-V4.1-Flash (552B MoE, MIT license): MiMo-V2.6-Pro leads on agentic benchmarks; DeepSeek retains advantages on pure reasoning and Chinese-language tasks
- vs. Qwen3.8 Max (2.4T MoE, proprietary API): MiMo-V2.6-Pro is fully open-weight; Qwen3.8 Max remains API-only, giving Xiaomi a significant edge for developers who want to self-host or fine-tune
- vs. Kimi K3 (2.8T, closed): Kimi K3 is larger and closed; MiMo-V2.6-Pro offers comparable composite scores with full weight access
The MIT license is the decisive differentiator. In a landscape where Alibaba's most capable models remain behind API walls and DeepSeek's open releases carry custom licensing terms for commercial use above certain revenue thresholds, Xiaomi's permissive approach is a genuine competitive advantage for the developer community.
The Anthropic Shadow
The release cannot be discussed without addressing the controversy that preceded it. On September 10, 2026 — twelve days before MiMo-V2.6-Pro shipped — Anthropic published a threat intelligence report identifying a campaign it tracked as GTG-16008. The report alleged that Xiaomi had routed over 400,000 user conversations from its MiMo chatbot through third-party tools called "OpenClaw" and "OpenCode" to Claude, between March and April 2026, without user consent, with the apparent intent of extracting Claude's capabilities to improve Xiaomi's own models.
Xiaomi has not issued a public response to these specific allegations. The company's silence has been interpreted variously as legal caution, strategic indifference, or an implicit acknowledgement that the practice — widely referred to in the industry as a "distillation attack" — is difficult to defend publicly even if it is technically difficult to prove.
What makes the situation genuinely complicated is that there is no direct public evidence linking the outputs of the GTG-16008 campaign to the final weights of MiMo-V2.6-Pro. The model's training documentation emphasises the RL and MOPD stages, and the $2.62 million training cost figure is consistent with a legitimate RL run at this scale. The Codersera technical guide↗ notes that the model's architecture and training methodology are independently coherent and do not require distillation to explain the performance gains.
Nevertheless, the provenance question is not going away. Enterprise legal teams evaluating MiMo-V2.6-Pro for production deployment will need to assess whether the Anthropic allegations — unresolved and publicly uncontested — create acceptable risk. For individual developers and researchers, the MIT license and the model's demonstrated capabilities are likely to outweigh those concerns.
Practical Access: How to Use MiMo-V2.6-Pro Today
For developers ready to experiment, access is straightforward:
- Hugging Face (open weights): Download the MiMo-V2.6-Pro-RL checkpoint↗ directly — no approval required. The MOPD checkpoint, which addresses tool-call repetition and adds open-domain capabilities, is available at XiaomiMiMo/MiMo-V2.6-Pro-MOPD↗
- Xiaomi API: OpenAI- and Anthropic-compatible endpoints via mimo.mi.com↗, priced at $0.435/M input tokens and $0.87/M output tokens; cached input at $0.0036/M
- OpenRouter: Available for developers who prefer a unified API gateway
- MiMo Studio and MiMo Code: Xiaomi's own IDE integrations for coding workflows
- MiMo Desktop: An invite-only agentic application providing computer-use capabilities — GUI interaction, file management, cross-application workflows
Self-hosting requires significant infrastructure: the mxfp4-quantized weights occupy approximately 566 GB, and efficient inference typically requires an 8× H200 or 8× B200 GPU configuration with tensor parallelism enabled via vLLM or SGLang.
What Qwen 4 Means for This Moment
The timing of MiMo-V2.6-Pro's release is not incidental. At Alibaba's Apsara Conference on September 22, 2026 — the same day Xiaomi shipped MiMo-V2.6-Pro — Alibaba confirmed that Qwen 4 is currently in training, with no release date provided. The next-generation Qwen family is expected to scale to between 5 and 10 trillion parameters in future iterations, and Alibaba previewed an architectural foundation through the Qwen3.8-Flash-Next experimental release in August.
In the window between Qwen3.8 Max and Qwen 4's eventual arrival, MiMo-V2.6-Pro occupies a strategically important position: it is the most capable open-weight Chinese model available today, and it is accessible to anyone with a Hugging Face account. For the global developer community that has been building on DeepSeek's open releases, Xiaomi has just offered a compelling alternative — one that leads on automation benchmarks and costs a fraction of proprietary alternatives.
Whether the Anthropic controversy ultimately affects adoption will depend on how the legal situation develops and whether Xiaomi chooses to respond publicly. For now, the model's performance speaks loudly enough that most developers appear willing to proceed.
Links & Resources
- MiMo-V2.6-Pro official model page↗
- MiMo-V2.6-Pro-RL weights on Hugging Face↗
- MiMo-V2.6-Pro-MOPD checkpoint on Hugging Face↗
- VentureBeat: MiMo-V2.6-Pro debuts as top open-weights model↗
- Forkast News: MiMo-V2.6 ships open weights at frontier-class performance↗
- BenchLM Chinese AI model leaderboard↗
- The Decoder: Xiaomi's affordable flagship leads open models amid Anthropic claims↗
- Spheron Network: Deploy MiMo-V2.6-Pro on GPU cloud↗
Links & Resources
External links — opens in a new tab

🇨🇦 China Desk Correspondent · Toronto, Canada
Bridges the East–West gap — what China’s models mean for everyone else.

A Treatise on Functional Analysis
by Richard Murdoch Montgomery
Structures, dualities, and spectra — Banach spaces, Hilbert spaces, operator theory, and spectral decompositions for the working mathematician.

Artificial Intelligence: Origins and Developments
by Richard Murdoch Montgomery
A comprehensive survey of AI from Turing machines to deep learning — neural networks, expert systems, and the philosophical debates that shaped the field.

CM1 Complete Study Material: Actuarial Mathematics
by Richard Murdoch Montgomery
The comprehensive guide for the CM1 actuarial exam — compound interest, annuities, life tables, reserving, and profit testing.

Medical AI
by Richard Murdoch Montgomery
Machine learning in clinical medicine — diagnostic imaging, drug discovery, electronic health records, and the ethics of algorithmic care.
Comments
Open discussion — no account needed. Be respectful.
More from Chinese Models Desk
The Claude Distillation Dispute: Seven Chinese AI Labs, Nearly 200 Million Exchanges and a Contested Model Release
Anthropic alleges that seven China-based AI labs conducted industrial-scale, unauthorized campaigns to extract capabilities from Claude, while Beijing rejects the accusations as groundless. The evidence also raises privacy, model-security and open-weights questions that the Chinese AI ecosystem will need to address.
Wei LianZhipu's GLM-5.3-FlashX Hits 200 Tokens Per Second — and the AI That Built It Is Running on Chinese Chips
Z.ai's new FlashX serving tier for GLM-5.3-Flash achieves a fourfold speed leap over the standard tier by deploying an AI agent to rewrite its own inference stack — all on a 100,000-accelerator cluster of domestically produced Chinese chips. It's the most concrete demonstration yet that China's sovereign compute ambitions are becoming real.
Sophia ChenDeepSeek-V4.1-Flash Rewrites the Rules on Agent Memory: 552B Parameters, MIT License, and a KV Cache That Fits in Your Pocket
DeepSeek's latest model slashes KV cache memory to 890 bytes per token through a radical Causal Encoder-Decoder architecture — then releases the weights for free under MIT. The result is a multimodal 552B-parameter model that outperforms its own V4-Pro flagship on agentic benchmarks at a fraction of the cost.
Wei Lian