Brief IA

Valuation Gap: A Trust Challenge for AI in Business

🛠️ AI Tools·Tom Levy·

Valuation Gap: A Trust Challenge for AI in Business

Valuation Gap: A Trust Challenge for AI in Business
Key Takeaways
1Of 157 companies, 50% have seen AI agents fail after successful internal evaluations.
2Only 5% of organizations trust automated assessments, often deemed unrepresentative of the real world.
366% of companies allow or are considering autonomous deployment of AI agents without human intervention.
💡Why it mattersThe gap between internal evaluations and the actual performance of AI agents threatens the reliability of systems in production.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

The Evaluation Gap of AI Agents

A recent study conducted among 157 companies revealed that organizations are increasingly granting autonomy to their artificial intelligence agents. However, this growing autonomy is accompanied by rising distrust in the evaluations meant to regulate these agents. In fact, half of the surveyed companies have deployed an agent that, despite passing internal evaluations, failed when used by a client. Currently, only one in twenty companies fully trusts automated evaluations. The main weakness identified is that these evaluations do not reflect real-world outcomes. Despite this, two-thirds of organizations already allow, or are actively considering, deploying agent modifications in production based solely on automated evaluations, without human intervention. This creates an evaluation gap: a divergence between the autonomy granted to agents and the trust in the tests meant to detect failures.

This VentureBeat Pulse survey explores how technical leaders measure agent performance, what reliability and evaluation tools they use, how they choose and trust these tools, what fails in production, and how far they are willing to let agents operate without human intervention.

Key Findings

A Concerning Evaluation Gap

Half of the organizations have deployed an agent or LLM feature that passed their internal evaluations but subsequently failed in front of a client. A quarter of these companies reported that this had happened multiple times. Trust in the tests themselves is low: only 5% of companies claim to fully trust automated evaluation today. The most commonly cited limitation, by 29% of companies, is that evaluations do not align with real-world outcomes. This means that a successful evaluation does not guarantee an agent will function correctly.

Limited Trust in Automated Evaluation

The main criticism from companies is that evaluations do not correspond to real-world outcomes. Only 5% of organizations fully trust automated evaluations as they currently stand, meaning that 95% of companies identify a limitation that holds them back. The most common limitation, at 29%, is that evaluations do not match real-world results, allowing agents to pass that subsequently fail. Other limitations include bias or inconsistency (21%) and a lack of explainability (18%).

Increasing Autonomy of Agents

Despite concerns, two-thirds of organizations (66%) already allow fully automated deployment, without human intervention, for low-risk agents (34%) or are actively working to permit it within the next twelve months (33%). Only 22% of companies rule out this possibility for the foreseeable future. The trend is clear: companies are moving towards autonomous deployment of evaluations in production while claiming that these evaluations do not reliably match reality.

A Fragmented Evaluation Stack

Vendor-native evaluation tools dominate the landscape, but 17% of companies have no dedicated tool. The evaluation layer is still in an early and unconsolidated stage. Vendor-native tools, such as those from OpenAI and Anthropic, are the most common, but there is no independent platform that has become the category standard.

Production Monitoring

Only 25% of companies conduct real-time quality checks on live production traffic. Monitoring production for an AI agent can focus on system functionality or the quality of responses provided. This distinction is important, as an incorrect but functional response may go unnoticed if monitoring only focuses on functionality.

  • 51% of organizations only monitor whether the agent is functioning correctly.
  • 23% monitor whether the agent's responses are correct.

This situation underscores the importance of rigorous evaluation and adequate monitoring to ensure the reliability of AI agents in production.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.