AI Evaluation Gap Puts Enterprises at Risk
A significant disconnect between how much autonomy organizations are giving to AI agents and how confident they are in the evaluations that govern them is creating operational risks.
According to new research across 157 enterprises, half have already deployed an agent or LLM feature that passed internal testing but then failed with customers. Only 5% fully trust automated evaluation today, with a common complaint being that tests don’t align with real-world outcomes.
Despite these concerns, two-thirds of organizations either allow or are actively engineering toward deploying agents without human review based solely on automated evaluations—a trend where autonomy is outpacing assurance.
The research highlights:
- 50% have deployed an agent that passed evaluation but then failed with customers
- Only 5% fully trust automated AI evaluations
- The most-cited limitation: evaluations don’t align with real-world outcomes (29%)
- 66% already permit or are working toward fully automated deployments
The study also found that evaluation practices are fragmented, with many relying on basic tools from model providers rather than comprehensive platforms.
As AI adoption accelerates, organizations need to address this evaluation gap by implementing more robust testing frameworks and incorporating real-world feedback loops into their development processes. The current trajectory risks operational failures as enterprises grant increasing autonomy without commensurate assurance.