Brief IA

Anthropic: Claude Generates 80% of Code, a Major Turning Point

🤖 Models & LLM·Tom Levy·

Anthropic: Claude Generates 80% of Code, a Major Turning Point

Anthropic: Claude Generates 80% of Code, a Major Turning Point
Key Takeaways
1Claude, Anthropic's AI, now generates over 80% of the code, compared to less than 5% before February 2025.
2In April, Claude solved 97% of a security issue in 800 hours, outperforming human researchers.
3Since early 2026, one-third of bugs are detected before production thanks to automated reviews by Claude.
💡Why it mattersThe increasing automation by AI is redefining the role of engineers and accelerating technological development.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

In a memo recently published by Anthropic, it was revealed that Claude, their artificial intelligence system, is now responsible for writing over 80% of the company's code, a significant increase from less than 5% before February 2025. This evolution has allowed engineers to produce eight times more code than in 2024. According to Marina Favaro and Jack Clark, co-authors of the document, an AI system capable of designing its own successor could emerge sooner than most institutions expect.

Today, when an engineer at Anthropic needs to solve a problem, they no longer turn to a traditional code editor. Instead, they describe their needs to Claude, which generates, tests, and proposes solutions. In just eighteen months, the share of code generated by Claude has risen from a few percent to over 80%. Each engineer now oversees eight times more code than they produced themselves in 2024. Claude is even capable of correcting its own errors and, in some cases, surpassing human researchers in experimental decision-making.

Last April, Claude agents solved an open research problem in AI safety in a cumulative 800 hours, achieving a 97% success rate, while two human researchers had only closed 23% of the gap in one week. This performance illustrates Claude's ability to optimize the research process and exceed human limits.

Claude in the Production Loop

Before February 2025, Anthropic engineers copied and pasted code snippets suggested by Claude into their editor. Today, they submit a goal to Claude and receive code without human intervention between steps. By May 2026, Claude achieved a success rate of 76% on the most open tasks, up from 26% six months earlier.

Tens of thousands of training jobs can be disabled simultaneously, with a few lines of text sent to Claude, and two hours later, an engineer validates the fix. For a human working alone, this would have taken two to three days. Since early 2026, every modification submitted to Anthropic's codebase goes through automated review. This retrospective analysis showed that a third of the bugs responsible for past incidents on Claude could have been detected before reaching production.

The success of a session is determined by a Claude judge; a session is considered successful if the Claude Code agent has clearly accomplished the user's tasks without requiring corrections. Variations in workload can lead to short-term fluctuations in success rates.

From Execution to Scientific Judgment

As demonstrated last April, Anthropic published the results of a research project entirely conducted by Claude agents on an open AI safety problem. Two human researchers, in one week, had closed 23% of the gap between a weak supervisor and a strong model trained on correct responses. Over 800 cumulative hours and approximately €16,300 in computing costs, Claude achieved a 97% success rate. Humans had chosen the problem and defined the criteria, but for the rest, the agents designed each experiment themselves.

Regarding the ability to optimize training code, by May 2025, Claude Opus 4 multiplied the speed of the starting code by 3. Recently, Mythos Preview multiplied this speed by 52. A skilled human researcher tops out at a factor of 4, in four to eight hours of work.

Beyond the time factor, these results have a significant impact on work organization at Anthropic. Claude generates code faster than engineers can review it, and human review now slows down the entire chain. Agents have autonomously corrected over 800 API errors, reducing one class of incidents by a factor of a thousand. According to one engineer, a human would have needed four years to accomplish this work.

However, as Marina Favaro and Jack Clark suggest, the trajectory could still diverge. The capabilities of the models could plateau, and current tools might spread through the economy without crossing new thresholds. Humans could also retain control over research directions while execution becomes increasingly automated. Alternatively, an AI system could train its own successor, and the pace of progress would depend solely on the available computing power.

In the meantime, with Project Glasswing, Mythos Preview detected over ten thousand critical vulnerabilities in just a few weeks, including a 27-year-old bug in OpenBSD.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.