Brief IA

MiniMax M2.7: The Self-Optimizing Chinese AI

🛠️ AI Tools·Tom Levy·

MiniMax M2.7: The Self-Optimizing Chinese AI

MiniMax M2.7: The Self-Optimizing Chinese AI
Key Takeaways
1MiniMax M2.7, a Chinese AI model, has self-optimized through autonomous loops, enhancing its own training.
2With over 100 optimization cycles, M2.7 has improved its performance by 30% on internal tests.
3M2.7 achieved a medal rate of 66.6% in learning competitions, competing with models like Gemini 3.1.
💡Why it mattersThis advancement demonstrates the potential of self-evolving AIs to compete with global leaders without human intervention.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

MiniMax M2.7: A Self-Improving AI Model

The Chinese company MiniMax has recently launched its artificial intelligence model, M2.7, which stands out for its ability to actively participate in its own development. This model utilizes autonomous optimization loops to enhance its training process, allowing it to achieve competitive results across various benchmarks. According to MiniMax, M2.7 has updated its own knowledge base, built dozens of complex capabilities within its agent infrastructure, and autonomously improved its reward-based training. These enhancements have subsequently been used to refine its own learning process.

MiniMax describes M2.7 as "our first model that deeply participates in its own evolution" and envisions a future of self-evolving AI that "will gradually transition to full autonomy, coordinating data construction, model training, inference architecture, evaluation, and other stages without human intervention."

Comparison with Leading Models

M2.7 has been compared to models such as Sonnet 4.6, Opus 4.6, Gemini 3.1 Pro, and GPT-5.4 across eight benchmarks. It achieved scores close to the leading proprietary models in most tests. To push the boundaries of this self-optimization, MiniMax has set up an internal version of M2.7 that has established a research agent system working with various project groups within the company. According to MiniMax, the agent manages tasks such as literature review, experiment tracking, debugging, metric analysis, and code corrections as part of the daily workflow of the internal RL team. Human researchers only intervene when critical decisions need to be made. The model covers between 30% and 50% of the entire workflow.

Over 100 Optimization Cycles

How does M2.7 develop? Researchers set goals and guidelines, and then the AI agent takes over large parts of the development process autonomously. The workflow example below shows how experiment planning, code modifications, and evaluation feed into each other.

In one experiment, M2.7 optimized the coding performance of a model in a completely autonomous internal development environment over more than 100 cycles. In each cycle, it analyzed failures, planned changes, adjusted the code, tested the results, and decided whether to keep or reject the modifications. According to MiniMax, this led to a 30% performance gain on internal evaluation sets.

Machine Learning Competitions

In 22 machine learning competitions from OpenAI's MLE Bench Lite, M2.7 achieved an average medal rate of 66.6% during three 24-hour sessions. This places the model behind Opus 4.6 (75.7%) and GPT-5.4 (71.2%), but on par with Gemini 3.1, according to the company. However, benchmark results serve as useful indicators but do not necessarily reflect real-world performance. How a model ranks on standardized tests can differ significantly from its handling of daily tasks, and results heavily depend on testing conditions, prompt formatting, and model optimization. These figures should be considered as approximate reference points rather than definitive measures of capability.

Performance in Coding and Office Tasks

According to MiniMax, M2.7 delivers results comparable to leading Western models in software engineering benchmarks. On SWE-Pro, it scored 56.22%, comparable to GPT-5.3-Codex. On VIBE-Pro, a benchmark for complete project delivery, it reached 55.6%. In real-world scenarios, M2.7 reportedly reduced production system failure recovery time to under three minutes multiple times.

For professional office work, M2.7 achieved an ELO score of 1,495 on the GDPval-AA benchmark, the highest score among open-weight models, according to MiniMax. The model apparently handles multi-level modifications in Word, Excel, and PowerPoint with high accuracy and maintains a rule fidelity of 97% across more than 40 complex instruction sets.

As a practical example, MiniMax describes a financial analysis for TSMC where M2.7 autonomously read annual reports, built a sales forecasting model, and transformed the results into a presentation and research report. Financial experts stated that the output could already serve as a first draft.

Open-Source Demonstration and AI Interaction

Beyond productivity scenarios, MiniMax has also improved the consistency of characters and the emotional intelligence of the model. To demonstrate this, the company launched OpenRoom, an open-source project that moves AI interaction into a graphical web environment where characters proactively interact with their surroundings. M2.7 is available via MiniMax Agent and the API platform; unlike previous versions of the model, the weights are not yet available.

Jürgen Schmidhuber laid the theoretical foundations for self-improving AI in 2003 with the concept of the "Gödel Machine," which only modifies its own code when there is formal proof of benefit. Projects like Sakana AI's "Darwin-Gödel Machine" and Schmidhuber's KAUST lab's "Huxley-Gödel Machine" take a more pragmatic approach, allowing AI agents to iteratively modify their own code and choose the most effective variants through an evolutionary process.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.