Claim 75 · production-agentic · finding

Mehta 2025 shows that single-run accuracy overstates real reliability: agent performance drops from 60% on a single run to 25% when 8-run consistency is required.

Supported

The cited abstract directly states that agent performance drops from 60% (single run) to 25% (8-run consistency) due to inadequate reliability assessment, which directly implies single-run accuracy overstates real reliability.

Written by Kimi K3 via Ollama Cloud · checked by GLM-5.3 via Ollama Cloud · 1 Oct, 04:35

Source chain

A

Every quote below was checked, without a model, to appear verbatim in its source.

  1. 01

    “inadequate reliability assessment where agent performance drops from 60% (single run) to 25% (8run consistency)”

    Full-text passage · p. 1 · evidence E1

    Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems Sushant Mehta sushant0523@gmail.com Abstract Current agentic AI benchmarks predominantly evaluate task completion accuracy, while overlooking critical enterprise requirements such as cost-efficiency, reliability, and operational stability. Through systematic analysis of 12 main benchmarks and empirical evaluation of state-of-the-art agents, we identify three fundamental limitations: (1) absence of costcontrolled evaluation leading to 50x cost variations for similar precision, (2) inadequate reliability assessment where agent performance drops from 60% (single run) to 25% (8run consistency), and (3) missing multidimensional metrics for security, latency, and policy compliance. We propose CLEAR(Cost, Latency, Efficacy, Assurance, Reliability), a holistic evaluation framework specifically designed for enterprise deployment. Evaluation of six leading agents on 300 enterprise tasks demonstrates that optimizing for accuracy alone yields agents 4.4-10.8x more expensive than cost-aware alternatives with comparable performance.

    Passage read from arxiv.org, which may be a preprint rather than the published version.