New MiniMax M2.7 proprietary AI model is 'self-evolving' and can perform 30-50% of reinforcement learning research workflow

MiniMax has unveiled M2.7, a proprietary LLM capable of managing its own reinforcement learning harnesses and handling up to 50% of its development workflow. This release signals a strategic shift for the Chinese startup toward proprietary models to compete with global leaders like OpenAI and Google.
In the last few years, Chinese AI startup MiniMax has become one of the most exciting in the crowded global AI marketplace, carving out a reputation for delivering frontier-level large language models (LLMs) with open source licenses and before that, high-quality AI video generation models (Hailuo).
The release of MiniMax M2.7 today — a new proprietary LLM designed to perform well powering AI agents and as the backend to third-party harnesses and tools like Claude Code, Kilo Code and OpenClaw — marks yet a new milestone: Rather than relying solely on human-led fine-tuning, MiniMax has leveraged M2.7 to build, monitor, and optimize its own reinforcement learning harnesses.
This move toward recursive self-improvement signals a shift in the industry: a future where the models we use are as much the architects of their progress as they are the products of human research. The model is categorized as a reasoning-only text model that delivers intelligence comparable to other leading systems while maintaining significantly higher cost efficiency.
However, with M2.7 being proprietary for now, it is a sign once again that Chinese AI startups — for much of the last year, the standard-bearers in the world of the open source AI frontier, making them appealing for enterprises globally due to low (or no) costs and customization — are shifting strategy and pursuing more proprietary frontier models like U.S. leaders like OpenAI, Google, and Anthropic have been doing for years.
MiniMax becomes the second Chinese startup to release a proprietary cutting-edge LLM in recent months following z.ai with its GLM-5 Turbo, and rumors that Alibaba's Qwen team is also shifting to proprietary development in the wake of the departure of senior leadership and other researchers.
Technical achievement: The self-evolution loop
The defining characteristic of MiniMax M2.7 is its role in its own creation. According to company documentation, earlier versions of the model were used to build a research agent harness capable of managing data pipelines, training environments, and evaluation infrastructure.
By autonomously triggering log-reading, debugging, and metric analysis, M2.7 handled between 30 percent and 50 percent of its own development workflow.
This is not merely an automation of rote tasks; the model optimized its own programming performance by analyzing failure trajectories and planning code modifications over iterative loops of 100 rounds or more.
"We intentionally trained the model to be better at planning and at clarifying requirements with the user," explained MiniMax Head of Engineering Skyler Miao on the social network X. "Next step is a more complex user simulator to push this even further."
This capability extends to complex environments via the MLE Bench Lite, a series of machine learning competitions designed to test autonomous research skills. In these trials, M2.7 achieved a medal rate of 66.6 percent, a performance level that ties with Google's new Gemini 3.1 and approaches the current state-of-the-art benchmarks set by Anthropic's Claude Opus 4.6.
Performance evolution: MiniMax m2.7 vs. m2.5
When compared to its predecessor, M2.5, released in February 2026, the M2.7 model demonstrates significant gains in high-stakes software engineering and professional office tasks. Key performance metrics include:
- Software engineering: M2.7 scored 56.22 percent on the SWE-Pro benchmark, matching GPT-5.3-Codex.
- Hallucination reduction: M2.7 achieves a hallucination rate of 34 percent, lower than Claude Sonnet 4.6 (46%) and Gemini 3.1 Pro Preview (50%).
- System comprehension: On Terminal Bench 2, the model scored 57.0 percent, demonstrating deep operational logic.
- Skill adherence: On the MM Claw evaluation, M2.7 maintained a 97 percent adherence rate.
Access, pricing, and integration
MiniMax M2.7 is a proprietary model available through the MiniMax API. It maintains a cost-leading price point of 0.30 dollars per 1 million input tokens and 1.20 dollars per 1 million output tokens. MiniMax offers structured Token Plans ranging from a $10/month Starter tier to a $150/month Ultra-High-Speed tier. The model is already integrated into major developer tools including Claude Code and Cursor.
Source: VentureBeat
















