Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
Sushant Mehta
Why it has this license class
AChecked 30 Sept 2026. Open license (CC-BY, CC-BY-SA, CC0, public domain): full text indexed and used in synthesis.
| Source | License | Open-access status | Read as |
|---|---|---|---|
| openalex | cc-by | green | Green |
| arxiv | http://creativecommons.org/licenses/by/4.0/ | — | Green |
Abstract
BCurrent agentic AI benchmarks predominantly evaluate task completion accuracy, while overlooking critical enterprise requirements such as cost-efficiency, reliability, and operational stability. Through systematic analysis of 12 main benchmarks and empirical evaluation of state-of-the-art agents, we identify three fundamental limitations: (1) absence of cost-controlled evaluation leading to 50x cost variations for similar precision, (2) inadequate reliability assessment where agent performance drops from 60\% (single run) to 25\% (8-run consistency), and (3) missing multidimensional metrics for security, latency, and policy compliance. We propose \textbf{CLEAR} (Cost, Latency, Efficacy, Assurance, Reliability), a holistic evaluation framework specifically designed for enterprise deployment. Evaluation of six leading agents on 300 enterprise tasks demonstrates that optimizing for accuracy alone yields agents 4.4-10.8x more expensive than cost-aware alternatives with comparable performance. Expert evaluation (N=15) confirms that CLEAR better predicts production success (correlation $ρ=0.83$) compared to accuracy-only evaluation ($ρ=0.41$).
Claims built on this paper
D- Mehta 2025 argues that current agentic AI benchmarks predominantly measure task-completion accuracy while overlooking the requirements that matter for enterprise deployment, such as cost-efficiency, reliability, and operational stability. Trace →
- Mehta 2025 shows that single-run accuracy overstates real reliability: agent performance drops from 60% on a single run to 25% when 8-run consistency is required. Trace →
- To close these gaps, Mehta 2025 proposes the CLEAR framework (Cost, Latency, Efficacy, Assurance, Reliability), and reports from an evaluation of six leading agents on 300 enterprise tasks that optimizing for accuracy alone yields agents 4.4–10.8x more expensive than cost-aware alternatives with comparable performance. Trace →
- The evaluation gap that Mehta 2025 quantifies in benchmarks—missing multidimensional metrics and inadequate reliability assessment—is illustrated concretely by Patra et al. 2026, whose governance and abstention behavior was validated only in a controlled environment rather than under real-world deployment conditions. Trace →