Anthropic: Claude Enhances Code, Calls for AI Pause

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Anthropic recently published a detailed report through its Anthropic Institute, revealing that Claude, its artificial intelligence model, is now responsible for over 80% of the company's production code. This change marks a significant transformation in how Anthropic develops its software, with a notable acceleration in engineer productivity.
Since the launch of Claude Code in February 2025, Claude's contribution to production code has surged from modest figures to over 80%, and even 90% when including scripts and experimental code. One employee even shared that they hadn't written any code themselves in five months.
Despite this impressive increase, Anthropic acknowledges that lines of code are not a perfect indicator of productivity. An internal survey conducted in March 2026 among 130 employees revealed that the increase in output due to Mythos Preview is estimated to be around four times, although this figure is likely overstated. Anthropic highlights recent research from METR showing that developers tend to overestimate the productivity gains from AI.
In terms of quality, the code generated by Claude has reached a level comparable to that of human developers by the end of 2025 and is expected to surpass this level in the coming year. Claude has also demonstrated its effectiveness by detecting one-third of bugs before they reach production and providing over 800 fixes in April 2026, reducing a class of API errors by a factor of 1,000. According to one engineer, a human would have needed four years to accomplish this work.
Claude does not just produce code. Its research capabilities have also improved, with an increasing success rate on complex tasks. Since early 2026, the duration of tasks that Claude can handle autonomously has doubled every four months. In March 2024, Claude Opus 3 could manage tasks lasting four minutes, while a year later, Claude Sonnet 3.7 managed an hour and a half. Claude Opus 4.6 is now tackling tasks of 12 hours.
METR found that Claude Mythos Preview could work for "at least" 16 hours and was "at the upper end of what [METR] can measure without new tasks." If this trend continues, tasks lasting a day could be within reach this year, followed by tasks lasting a week in 2027, according to Anthropic.
Beyond mere code production, Anthropic is showing progress in AI-assisted research. In an internal optimization test where Claude had to make training code as fast as possible, Claude Opus 4 achieved an average speed of about 3x in May 2025. A year later, Mythos Preview reached around 52x. An experienced human researcher would need four to eight hours to achieve 4x.
In an analysis of actual research sessions at Anthropic, the company examined 129 moments where human developers took suboptimal detours. Claude Mythos Preview suggested the best next step in 64% of these cases, compared to 51% for Claude Opus 4.5 six months earlier. Anthropic calls this "an early signal that AI systems are improving in decision-making that AI research depends on."
The critical bottleneck, according to Anthropic, is what the company refers to as "research taste": the ability to choose the right problems and spot dead ends early. "The comparative advantage of humans, for now, still lies in the ability to see the big picture and think beyond the immediate task's boundaries," said an employee.
Whether this leap is achievable with current methods remains an open question. "It is really uncertain whether today's training methods and architectures could unlock this capability," the report states.
Anthropic puts this in perspective: paradigm shifts like the Transformer architecture occur with years in between. Most progress between these shifts is incremental work, exactly the type of workflow that Claude now manages well. Drawing from Edison’s famous quote about genius being one percent inspiration and ninety-nine percent perspiration, Anthropic writes: "We see the perspiration becoming increasingly automated."
Even if Claude never develops a good research taste, a conservative reading of the data implies a cumulative acceleration: each engineer contributes much more work than before, as humans only manage a single-digit percentage of directional decisions.
Anthropic outlines three future scenarios. In the first, the trend stagnates. Perhaps exponential curves turn out to be S-curves, or energy and chip bottlenecks slow things down. Anthropic considers this unlikely, as no flattening is visible so far.
In the second, efficiency gains continue but humans retain directional control. Companies of 100 people could accomplish the work of 10,000 or 100,000. Anthropic sees itself on this path but warns of risks such as authoritarian oversight and targeted manipulation campaigns. Amdahl's law also comes into play: at Anthropic, human code review has already become the new bottleneck.
The third scenario describes complete recursive autonomous improvement, where the pace of progress is limited only by computing power. The question of whether the alignment problem can be solved in this case is "something we are least certain about." Rare cases of misalignment could accumulate, "becoming more frequent but less understood until we lose control."
The most striking passage of the report concerns Anthropic's stance on slowing down AI development. "We believe it would be good for the world to have the option to slow down or temporarily suspend the development of cutting-edge AI to allow societal structures and alignment research to keep pace with technological advancement," the report states.
Anthropic claims it would slow down or suspend its activities if other leading developers did the same in a verifiable manner. The Anthropic Institute plans to research and build systems that would make a credible pause possible—mechanisms allowing leading developers to verify that others have indeed stopped or slowed down.
However, the obstacles are enormous. Training sessions are much easier to hide than missile silos. Their inputs are versatile, and the incentive to continue secretly is massive: "Anyone who continues while others take a pause could inherit the lead." A comparison to the INF treaty on nuclear weapons seems obvious, but these verification regimes took decades to build. "We do not have that time," writes Anthropic.
A unilateral pause by a single lab would be easy to implement immediately but would have much less impact. It would only change the leader without creating the broader deliberative process that is lacking. In the coming months, Anthropic plans to organize discussions involving policymakers, researchers, civil society, and other AI companies.
The debate around an AI pause has already been raised years ago, but these calls have not gained traction. In hindsight, this push seemed premature given what these systems could actually do. Whether this is fear-based marketing, as critics have already accused with Mythos, will likely only become clear with time.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.