⚡
Brief IA
›

Multi-Agent AI: Increased Speed, Slightly Improved Quality

🛠️ AI Tools·Tom Levy·

Multi-Agent AI: Increased Speed, Slightly Improved Quality

Multi-Agent AI: Increased Speed, Slightly Improved Quality
⚡
Key Takeaways
1AI agent teams cost between 1.8 and 5.1 times more than individual agents according to Vals AI
2Only one tested configuration shows a significant gain of 7.3 points, with most providing no real advantage
3Speed gains exist but come with an increase in token consumption and degraded coordination
4The effectiveness of teams heavily depends on the task, and returns diminish rapidly with the number of agents
💡Why it matters — These results call into question the relevance of using AI agent teams to improve quality, as the resource overhead is rarely offset by measurable benefits.

Trials conducted by Vals AI, Anthropic, and members of OpenAI converge on the same conclusion: agent teams are significantly more expensive and yield little improvement in results. Only one configuration tested by Vals AI shows a significant gain of 7.3 points, while most show none. Speed gains exist, but they come with a token cost and a "coordination tax" mentioned by OpenAI.

OpenAI: Time Gains Proportional to Cost

Noam Brown, a researcher at OpenAI, indicates that multi-agent systems primarily accelerate task execution without notable improvements in quality. He specifies that four agents complete a task twice as fast, but at double the cost, and this trend continues with 16 agents, showing slightly reduced efficiency. Brown emphasizes that the use of very large groups of agents remains underexplored due to prohibitively high costs. Eric Provencher, a developer at OpenAI, warns against the massive deployment of agents, arguing that coordination deteriorates and represents a "coordination tax" that could render operations unprofitable.

Vals AI Benchmarks: A Cost Overhead of 1.8x to 5.1x for Little Gain

Vals AI evaluated GPT-6 Sol and Claude Opus 5.5 on the Vibe Code Bench, in both individual and team configurations, at two reasoning levels: medium and maximum. Agent teams cost between 1.8 and 5.1 times more than individual agents. Out of four comparisons, only one showed a statistically significant improvement: the GPT-6 Sol team at medium reasoning, with a gain of 7.3 points. For the maximum reasoning level, no concrete advantage was observed for Sol or Opus 5.5 when grouped in a team. According to Vals AI, the additional cost associated with agent teams is not justified in most cases, especially when models are already reaching their maximum capacity. The benchmark graph correlates cost per application with the score achieved, showing that teams lead to a marked increase in costs for very limited score improvements.

⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

Anthropic: Speed Gains, Diminishing Returns, and Task Dependency

Anthropic found that quality gains diminish as the number of agents increases during tests with Opus 5.5. Larger teams reach a given performance level more quickly, but moving from ten to one hundred agents only slightly increased scores after 24 hours. In separate trials on ProgramBench, the acceleration was accompanied by a rise in token consumption. Noam Brown specifies that the effect of teams heavily depends on the task: web searching and mathematics lend themselves well to parallelization, while writing a novel does not. He illustrates this point by explaining that deploying 10,000 agents on a novel would be as ineffective as assigning 10,000 people to it.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.