Frontier Models
Xiaomi MiMo-V2.6 Pro and Flash Match Closed Frontier Models on Agent Benchmarks
The open-weight omnimodal models incorporate scaled reinforcement learning to advance performance in coding and spatial reasoning while releasing the complete training stack for community use.
MiMo-V2.6 is a series of natively omnimodal open-weight models from Xiaomi focused on scaling reinforcement learning for enhanced coding, 3D reasoning, and agent capabilities.
On September 22, 2026, Xiaomi introduced the MiMo-V2.6 series consisting of two primary omnimodal open-weight models known as MiMo-V2.6-Pro and MiMo-V2.6-Flash. These models represent an effort to scale reinforcement learning on verifiable tasks to push the boundaries of open-source AI capabilities in areas that have traditionally favored closed systems. The release includes not only the model weights but also the complete technical report, training environments, and RL code, all made available under a permissive license on Hugging Face. The Pro model in particular has achieved scores that place it among the top open models available today. This release includes extensive resources that allow others to reproduce and extend the work. The focus on verifiable tasks in the RL process enables the models to improve through self-directed exploration and feedback mechanisms.
What background informs the development of the MiMo-V2.6 series?
The development of MiMo-V2.6 comes amid increasing competition in the open-source AI space where companies seek to match the performance of proprietary systems from leading laboratories. Previous iterations of open models have shown progress in general capabilities but often lagged in specialized areas like agentic reasoning and multimodal integration. Xiaomi's approach focuses on reinforcement learning to enable continuous capability expansion through exploration and feedback loops that refine decision making over multiple iterations. The landscape of frontier models has seen closed systems lead in agentic tasks, but open efforts have been closing gaps through targeted scaling of compute on complex verifiable problems. Open-sourcing the full training stack provides the community with resources to build upon the work including the technical report that details the methods used and the code for RL training. Such transparency contrasts with some closed models where details remain proprietary and unavailable for external scrutiny or adaptation.
The timing aligns with broader industry trends toward omnimodal systems that handle text, images, and potentially other inputs seamlessly. Gains in coding and 3D spatial reasoning allow the models to generate interactive 3D scenes and objects in Blender from text or image prompts. This capability opens new avenues for applications in game development and design workflows where spatial understanding is essential. The models demonstrate particular strengths in areas that have challenged open systems previously. These include advanced coding tasks and 3D reasoning that enable practical applications in creative and technical domains. Community access to these tools could accelerate experimentation beyond what was possible with earlier open releases.
What distinguishes the Pro and Flash variants in the series?
MiMo-V2.6-Pro serves as the most powerful flagship reasoning model with full omni-modal capabilities suited for demanding applications. In contrast, MiMo-V2.6-Flash offers high intelligence at lower cost, making it suitable for a wider range of applications where efficiency and reduced computational demands matter most. Both completed 30 RL steps each in under six days over approximately 750k trajectories. Additional variants include MiMo-V2.6-Pro-UltraSpeed and MiMo-V2.6-Distill-Qwen-9B, which provide options for different performance and size trade-offs depending on deployment needs. The Pro model emphasizes maximum capability while Flash prioritizes accessibility in terms of computational requirements during inference and fine-tuning. The models demonstrate gains in coding and 3D spatial reasoning that support generation of interactive 3D scenes and objects in Blender from prompts.
The distinction between variants allows users to select based on specific requirements for intelligence versus cost. Pro targets scenarios requiring top performance on complex agent benchmarks while Flash addresses use cases where lower latency and resource use provide advantages. This dual approach broadens the potential adoption across different segments of the AI ecosystem. The open weights enable customization for specialized domains through further training on domain-specific data. Such flexibility supports integration into existing workflows without the constraints often associated with closed API services.
How do benchmark results position these models against competitors?
This score places it ahead of other open-source offerings and indicates strong overall intelligence across evaluated dimensions. On agent benchmarks, the tying score with closed models suggests parity in certain interactive task evaluations that test reasoning and tool use. The ProgramBench result of 26.5 highlights improvements in programming-related agent performance compared to GPT-5.6 Sol. These metrics provide standardized ways to assess progress across different model types and allow direct comparisons that were previously difficult due to access limitations. The open nature allows independent verification of these claims through the released technical report and evaluation setups.
Comparisons with Kimi K3 and Qwen3.8 Max show the Pro model surpassing those open alternatives on the Intelligence Index. This positioning strengthens the case for open models in competitive evaluations where closed systems have historically dominated. The results reflect the effectiveness of the RL scaling strategy applied during training. Stakeholders can now reference these benchmarks when selecting models for agentic applications in coding and spatial domains.
What technical details characterize the training and architecture?
| Model Variant | Intelligence Index Score | Agents' Last Exam Score | ProgramBench Score | RL Training Cost |
|---|---|---|---|---|
| MiMo-V2.6-Pro | 46.32 | 31.6 | 26.5 | $2.62M |
| MiMo-V2.6-Flash | Not specified | Not specified | Not specified | $0.85M |
| Claude Opus 5 | Not specified | 31.6 | Not specified | Not specified |
| GPT-5.6 Sol | Not specified | Not specified | 25.0 | Not specified |
- Full technical report detailing methods and results
- Training environments for replication
- RL code for scaling reinforcement learning
- Model weights hosted on Hugging Face under permissive license
The RL training process involved running 30 steps for each model variant. This was accomplished in under six days using a large number of trajectories to refine the models' decision-making abilities across the evaluated tasks. By focusing on verifiable tasks, the approach allows the model to receive feedback that drives improvement without human intervention in every step. This method supports the goal of expanding the capability frontier through iterative self-improvement. The architecture supports native handling of multiple input types which contributes to the observed gains in integrated reasoning scenarios. The released code provides the foundation for others to experiment with similar scaling approaches on their own hardware setups.
Training environments released alongside the models include the necessary setups for running the RL processes described in the technical report. These resources lower the entry barrier for researchers interested in replicating or extending the work. The permissive license on the weights encourages broad experimentation and derivative model creation. Such openness facilitates rapid iteration within the open-source community compared to scenarios where only final weights are shared without supporting code.
What market and stakeholder implications follow from this release?
For developers and researchers, access to the full stack lowers barriers to experimenting with advanced RL techniques on complex tasks. This could accelerate innovation in open-source communities that previously lacked such resources for high-performance agent development. Enterprises may find the Flash variant attractive for cost-sensitive deployments while the Pro offers high-end performance for complex applications in coding and design. The omnimodal nature supports use cases in creative industries and software development where integrated reasoning across modalities provides value. The release challenges the notion that only closed models can achieve top-tier agent performance on standardized benchmarks.
Stakeholders in the AI ecosystem gain new reference points for evaluating open versus closed options in frontier capabilities. The availability of training code enables customization for specific industry needs without reliance on external providers. This shift may influence procurement decisions and investment priorities toward solutions that offer greater transparency and control. Broader adoption could emerge as teams integrate the models into production workflows for tasks requiring spatial and coding intelligence.
How have experts and the community reacted to the MiMo-V2.6 models?
Today, we are releasing and open-sourcing the MiMo-V2.6 series. This marks a key step in our exploration of the RSI path: scaling RL compute on verifiable, complex tasks, so the model can continuously expand its capability frontier through exploration and feedback.Xiaomi MiMo team
The statement from the team emphasizes the strategic direction toward recursive self-improvement through RL. Community discussions on platforms like Hugging Face likely focus on the practical applications of the open code and environments. Reactions highlight the significance of open-sourcing the RL code as it enables broader participation in advancing these techniques beyond what isolated research groups could achieve alone. This transparency fosters collaboration that could lead to further refinements and new applications in the months ahead.
The emphasis on verifiable tasks and feedback loops has drawn attention as a replicable method for capability expansion. Analysts note that the cost figures for training provide context on the resources required to reach these performance levels in open settings. Such details help calibrate expectations for future open releases aiming to compete on similar benchmarks.
What developments can be expected in the coming months?
Further iterations may build on the released resources to improve efficiency or add new modalities to the existing omnimodal foundation. The community might produce fine-tuned versions based on the Distill-Qwen-9B model for specific domains such as specialized coding environments or 3D asset creation pipelines. Continued scaling of RL compute could push scores even higher on the Intelligence Index as additional trajectories and steps are applied. Monitoring updates from Xiaomi will reveal the next steps in this exploration of the RSI path.
Independent researchers are positioned to test the models on additional benchmarks not covered in the initial release. This could uncover strengths or limitations that inform subsequent training runs. The permissive license supports the creation of derivative works that extend the core capabilities into new application areas. Overall the release sets a precedent for comprehensive open-sourcing that may influence how other organizations approach similar model launches.
Frequently asked
How does MiMo-V2.6-Pro perform compared to closed models on agent benchmarks?
MiMo-V2.6-Pro ties Claude Opus 5 at 31.6 on Agents' Last Exam and scores 26.5 on ProgramBench compared to GPT-5.6 Sol's 25.0. These results indicate competitive performance on tasks involving reasoning and programming agents.
What resources were open-sourced with the MiMo-V2.6 release?
The release includes the full technical report, training environments, RL code, and model weights on Hugging Face under a permissive license. These elements allow replication and extension of the work by the broader community.
What is the training cost associated with the MiMo-V2.6 models?
RL training for MiMo-V2.6-Flash cost $0.85 million while MiMo-V2.6-Pro required $2.62 million. Both variants completed 30 RL steps in under six days using around 750k trajectories.
Sources
- Xiaomi — MiMo-V2.6-Pro scores 46.32 on the Artificial Analysis Intelligence Index and the quote on releasing the series with full resources.
- Xiaomi — Announcement of the 2026-09-22 MiMo-V2.6 Series Release including Pro and Flash variants.
- Hugging Face — Release of MiMo-V2.6-Pro-RL, MiMo-V2.6-Flash-RL, and MiMo-V2.6-Distill-Qwen-9B models in the collection.