Authors: Stephan Rabanser, Sayash Kapoor, and Arvind Narayanan
Suppose you hear about a new AI agent that can boost productivity—shopping for you, writing code, sending emails, or handling customers on your behalf. Should you trust it? Can the agent complete tasks reliably enough? After all, there have already been plenty of horrorstories about agents going wrong, such as this incident.
Surprisingly, although the lack of reliability in AI agents has long been widely recognized, the AI industry currently has neither good tools for measuring reliability nor even a good definition of “reliability.”
Arvind and Sayash have been thinking about this problem for a long time. Last fall, postdoctoral researcher Stephan Rabanser joined us. His doctoral research focused on reliability in simpler, more traditional AI systems. We invited several independent researchers to join us and published the results of our work, which we hope will provide a comprehensive way to measure reliability. Our draft paper is titled “Towards a Science of AI Agent Reliability.”
We drew on research from many other fields, including nuclear and aviation safety, and decomposed reliability into 12 distinct dimensions. We evaluated 14 models using two complementary benchmarks and found that the rapid capability gains of the past two years have led to only limited improvements in reliability. See our interactive dashboard.
Although our findings are still preliminary, we hope they can help explain why many people in the industry feel confused: despite AI agents performing strongly on capability benchmarks, their economic impact has been emerging only gradually. [^1] To help the broader community track reliability systematically, we plan to launch an “AI Agent Reliability Index.” We hope this will encourage both researchers and industry to devote more effort to improving reliability.
