By Stephan Rabanser, Sayash Kapoor, and Arvind Narayanan
Imagine you hear about a new AI agent that can boost productivity by shopping, coding, sending emails, or handling customer service on your behalf. Would you trust it? Can this agent get the job done reliably enough? After all, many horror stories of agents going wrong are already well known.
Surprisingly, although the lack of reliability in AI agents has long been a well-known issue, the AI industry currently has neither good tools to measure reliability nor even a solid definition of "reliability" itself.
Arvind and Sayash have been thinking about this problem for a long time. Last fall, postdoctoral researcher Stephan Rabanser joined us; his PhD research focused on reliability issues in simpler, more traditional AI systems. We recruited several independent researchers and published what we hope is a comprehensive effort to measure reliability. Our draft paper is titled Towards a Science of AI Agent Reliability.
We drew insights from various fields such as nuclear safety and aviation safety, and broke reliability down into 12 distinct dimensions. We evaluated 14 models across two complementary benchmarks and found that the rapid gains in capability over the last two years have yielded only modest improvements in reliability. Our interactive dashboard can be found here.
Although our current findings are preliminary, we hope they will help explain a puzzle shared by many in the industry: why AI agents perform spectacularly on capability benchmarks, yet their economic impact has been slow to materialize.1 To help the community systematically track reliability, we plan to launch an AI Agent "Reliability Index." We hope this will encourage researchers and the industry to invest effort in improving reliability.
