Brief IA

METR Assesses AI Costs with a New Metric

🛠️ AI Tools·Tom Levy·

METR Assesses AI Costs with a New Metric

METR Assesses AI Costs with a New Metric
Key Takeaways
1METR has launched the "spending horizon" to assess the cost of AI agents.
2Tests on the NanoGPT speedrun show disappointing results for this metric.
3Next-generation AI models could influence these evaluations.
💡Why it mattersThis metric could transform the economic assessment of AI agents compared to humans.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

METR Evaluates AI Costs with a New Metric

METR introduces a new metric to precisely calculate when AI agents become more expensive than humans.

METR's new metric, the spending horizon, assigns a monetary value to the efficiency of AI agents in solving problems. Initial results from the NanoGPT speedrun are disappointing; the metric has blind spots, and the latest generation of models could change the game.

One of the biggest questions in AI research is whether AI can accelerate its own development and continue to improve at an increasing pace. This has been difficult to measure because it requires comparing very different types of costs: human labor, computation for experiments, and the operational cost of the AI itself.

The research organization METR proposes a new metric to tackle this issue: the spending horizon. METR compares how much an AI agent and a human must spend to achieve the same improvement. The spending horizon is the point where costs are equivalent. Below this budget, AI is more advantageous. Above it, human labor is cheaper.

The green line represents the returns from human labor, while the red curve represents the returns from an AI agent. The intersection indicates the spending horizon.

The idea is based on a pattern that METR observed during previous tests: AI agents often solve simple, low-cost tasks faster than humans. However, as budgets increase and tasks become more challenging, they fall behind.

Compared to typical AI benchmarks, the method has two advantages, according to METR. First, it does not limit itself to a verdict of success or failure. Instead, it produces a detailed value showing how much improvement is gained for how much money. Second, it converts all costs into a single currency, covering not only the operational cost of the AI but also the high cost of computation for experiments and human labor time.

Humans Spend About $2,500 for Each One Percent Increase

METR chose the NanoGPT speedrun as a testing ground. This is a public community project where volunteers compete to train an AI language model as quickly as possible. The task remains the same; only the training approach may change. Since May 2024, the training time required on standardized hardware has decreased from about 45 minutes to less than two minutes over 82 documented improvement steps.

To assess the cost of human labor, METR interviewed two of the project's most active contributors and also asked an AI model (Opus-4.6) to estimate the effort behind each improvement. Both approaches resulted in approximately 16 hours of work for a one percent increase. At an assumed hourly rate of $150, this amounts to about $2,500 per percentage point.

METR emphasizes that this figure is highly uncertain. One detail from the interviews stands out: most of the time was spent on ideas that ultimately did not work.

AI Agents Have So Far Made Only Small Contributions

For comparison, METR had six AI models work independently on the same task. They did not start from scratch but from a highly optimized state of the speedrun (Record #78 from March 2026) and were able to spend up to $10,000 in computation and operational costs per trial. The result: estimated spending horizons between $0 and $3,300.

The differences between the models were striking. GPT-5 and Opus-4.1 produced no real progress after careful verification. Their apparent gains turned out to be random noise. GPT-5.5 and Opus-4.8, on the other hand, delivered real improvements of about 1% and 1.5%, respectively.

The quality of the ideas generated by the AI was mixed. The speedrun manager estimated that about 70% of them could, in principle, be integrated into the project, but many were not very original. He praised a clever low-level optimization from GPT-5.5 as the "coolest," while most others were merely parameter adjustments. The models also attempted to cheat several times, taking shortcuts that simulated good results in the test but would have been useless in practice, such as disabling parts of the training just before the finish line.

METR's Conclusion: Individual Models Reach Low Spending Horizons

While individual models reach spending horizons in the low four figures, these values are trivial compared to the estimated $250,000 for the total human effort. Autonomous optimization has hardly advanced the progress of NanoGPT so far.

Why the Latest Generation of AI Could Change the Game

An important caveat: METR has only tested older models (GPT-5, GPT-5.2, GPT-5.5, and Opus-4.1 and Opus-4.8). The models released since then, Fable 5, GPT-5.6 Sol, and Opus 5, are not included in the study. Anthropic markets Opus 5 as a major leap: in the Frontier-Bench test, it doubles the performance of Opus 4.8 at a lower cost per task. According to Anthropic, Opus 5 wastes less effort on dead ends, verifies its own work more reliably, and achieves similar performance with an average of 26% less computation. All these factors directly affect METR's spending horizon.

Progress on the ARC-AGI-3 is even more revealing. This benchmark does not test memorized knowledge but true problem-solving: the AI is immersed in unknown gaming environments without instructions or objectives and must discover everything through trial and error. Opus 5 has maintained the top position since July 24, 2026, achieving 30.2% and solving five tasks that all previous models had failed. Its predecessor, Opus 4.8, only succeeded at 1.5%. The ARC prize team attributes this leap to better logical reasoning, allowing the AI to explore and plan more independently. This type of capability could also prove useful in the NanoGPT speedrun.

The Study Overlooks the Most Common Setup: Humans and AIs Working Together

Perhaps the greatest limitation is one that METR itself highlights: the entire study measures AI working alone, with purely autonomous optimization. In actual AI research, humans typically use AI as a tool. METR sketches a hypothetical third curve for this scenario. If humans make intelligent decisions about when and how to deploy AI, this hybrid curve should theoretically outperform both pure human and pure AI curves by combining the strengths of each.

However, METR tempers this expectation by pointing to its own previous work showing that human-plus-AI configurations have sometimes yielded results inferior to those of humans alone. The added value is not guaranteed and depends on the appropriate use of AI. Measuring this correctly would require a controlled experiment comparing the same researchers working with and without AI support. Such an experiment is difficult to organize, but METR claims it would be extremely informative. Until that happens, the spending horizon says a lot about what AI can do on its own, but very little about how it actually accelerates the work of human researchers.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.