Brief IA

OpenAI: GPT-5.6 Surpasses GPT-5.5 with Autonomous Advancement

🤖 Models & LLM·Tom Levy·

OpenAI: GPT-5.6 Surpasses GPT-5.5 with Autonomous Advancement

OpenAI: GPT-5.6 Surpasses GPT-5.5 with Autonomous Advancement
Key Takeaways
1OpenAI's GPT-5.6 Sol has refined the Luna model with a simple, minimally detailed prompt.
2In the RSI benchmark, Sol surpasses GPT-5.5 by 16.2 points, demonstrating a notable improvement.
3OpenAI envisions a near future where an "automated researcher" could become a reality.
💡Why it mattersThis advancement marks a significant step towards AI autonomy, reducing the need for human intervention in model improvement.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

OpenAI: GPT-5.6 Sol Surpasses GPT-5.5 with Autonomous Advancement

OpenAI's new AI model, GPT-5.6 Sol, is capable of autonomously optimizing smaller models. According to the company, a brief prompt was sufficient for Sol to independently identify training configurations, select GPUs, and execute the post-training script for the Luna model.

On a new internal index measuring recursive self-improvement (RSI), which evaluates a system's ability to evolve on its own, GPT-5.6 Sol scored 16.2 points higher than its predecessor, GPT-5.5.

Context of Autonomous Post-Training

On July 11, 2025, OpenAI employee Jason Liu clarified the context of autonomous post-training. Sol did not create a complete training recipe from scratch, as most configurations already existed thanks to Sol's post-training. The actual task was to adapt this configuration for the smaller Luna model and execute the training work. According to Liu, this would have otherwise "taken two researchers about two additional weeks, so it's a huge advancement."

Improvement of Research Capabilities

OpenAI claims that the new AI models accelerate the development of AI itself. The flagship model GPT-5.6 Sol autonomously post-trained the smaller Luna model. After Luna's initial pre-training, Sol optimized it for specific skills and behaviors autonomously.

A researcher provided Sol with a "fairly underspecified prompt" via the Codex platform. The instructions asked the model to find the right training configurations, choose appropriate GPUs, launch the training script, and verify that everything was functioning correctly.

Self-Improvement Benchmark

To measure these capabilities, OpenAI built an internal evaluation suite based on real-world AI research tasks. These tasks include:

  • Debugging search systems
  • Optimizing kernels and training recipes
  • Executing machine learning experiments
  • Improving another model

GPT-5.6 Sol scored 16.2 points higher than GPT-5.5 on the aggregated RSI index, placing Sol at the top of the model hierarchy, followed by the Terra and Luna variants, then GPT-5.5 and GPT-5.4.

Implications of Autonomous Improvement

Recursive self-improvement in AI research refers to an AI system's ability to enhance itself, where each cycle of gains makes the system even more capable of improving. This creates a feedback loop. This term is central to AI safety research, as a system capable of recursive improvement could, in theory, trigger a rapid explosion of capabilities.

Increased Use of GPT-5.6

OpenAI indicates that its researchers are using GPT-5.6 Sol throughout the development cycle, from debugging and optimizing training systems to executing experiments and analyzing results. Even during internal testing, the average daily token production per active researcher more than doubled compared to the previous record set by GPT-5.5. Pull requests and experiments per researcher have also increased, allowing teams to turn ideas into results more quickly.

The company's adoption figures over the past six months show significant growth:

  • The share of compute allocated to internal coding inference has increased 100 times.
  • Agent-based token usage has surged by about 22 times.

OpenAI acknowledges that these metrics do not directly measure research progress but asserts that they demonstrate how quickly AI-assisted work is evolving.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.