Production agentic workflows and evaluation
What makes an agentic workflow fit for daily operation: evaluation beyond task success, deployment practice, and the failure modes of autonomous agents.
Synthesised claims
AEach claim cites its evidence; a second model has checked it against that evidence.
- 01
Mehta 2025 argues that current agentic AI benchmarks predominantly measure task-completion accuracy while overlooking the requirements that matter for enterprise deployment, such as cost-efficiency, reliability, and operational stability.
- 02
Mehta 2025 shows that single-run accuracy overstates real reliability: agent performance drops from 60% on a single run to 25% when 8-run consistency is required.
- 03
To close these gaps, Mehta 2025 proposes the CLEAR framework (Cost, Latency, Efficacy, Assurance, Reliability), and reports from an evaluation of six leading agents on 300 enterprise tasks that optimizing for accuracy alone yields agents 4.4–10.8x more expensive than cost-aware alternatives with comparable performance.
- 04
Eranga et al. 2025 characterize production agentic workflows as dynamic pipelines of multiple specialized agents using different LLMs, tool-augmented capabilities, and orchestration logic, and they offer a structured engineering lifecycle plus nine core best practices—including tool-first design over MCP and single-responsibility agents—for building them.
- 05
Eranga et al. 2025 present their containerized, Kubernetes-orchestrated, MCP-accessible system with Responsible-AI mechanisms as a robust template that organizations can adapt across domains such as compliance automation, media generation, analytics, and enterprise RPA.
- 06
Patra et al. 2026 build governance directly into their autism-screening architecture: an orchestration layer coordinates specialized agents for consent verification, bias and applicability checks, and confidence evaluation, deliberately kept as separate responsibilities so that every decision remains transparent and auditable.
- 07
Patra et al. 2026 acknowledge that their system is not yet production-validated: its abstention and governance behavior was evaluated only in a fully controlled environment, and how the system would integrate into existing clinical workflows remains unevaluated.
- 08
Eranga et al. 2025 and Patra et al. 2026 independently converge on separation of concerns as the foundation of auditable production agentic systems: Eranga prescribes single-responsibility agents and clean separation between workflow logic and MCP servers, while Patra assigns each architectural layer a clear responsibility so governance checks stay visible and testable.
- 09
The evaluation gap that Mehta 2025 quantifies in benchmarks—missing multidimensional metrics and inadequate reliability assessment—is illustrated concretely by Patra et al. 2026, whose governance and abstention behavior was validated only in a controlled environment rather than under real-world deployment conditions.
- 10
MCP is emerging as shared infrastructure across the production-agentic literature: Eranga et al. 2025 prescribe tool-first design over MCP with clean separation between workflow logic and MCP servers, while Yue et al. 2026's perspective on MCP-native AI scientist ecosystems cites MCP construction tooling such as Code2MCP alongside work on measuring agents in production.
Anchor papers
B| Paper | Year | License |
|---|---|---|
| Agents of Chaos arXiv (Cornell University)· preprint From the reviewed export of the private instance (2026-09-30): identifier only. | 2026 | Orange |
| Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems arXiv (Cornell University)· preprint From the reviewed export of the private instance (2026-09-30): identifier only. | 2025 | Green |
| A Practical Guide for Designing, Developing, and Deploying Production-Grade Agentic AI Workflows arXiv (Cornell University)· preprint From the reviewed export of the private instance (2026-09-30): identifier only. | 2025 | Green |
| Orchestrating Human-AI Teams: The Manager Agent as a Unifying Research Challenge arXiv (Cornell University)· preprint From the reviewed export of the private instance (2026-09-30): identifier only. | 2025 | Orange |
| AFlow: Automating Agentic Workflow Generation arXiv (Cornell University)· preprint From the reviewed export of the private instance (2026-09-30): identifier only. | 2024 | Orange |
Found through citations
C| Paper | Year | License |
|---|---|---|
| Building MCP-native hierarchical AI scientist ecosystems: a perspective on scaling multi-agent scientific discovery Frontiers in Artificial Intelligence discovered: cites 1 anchor(s) (AFlow: Automating Agentic Workflow Generation); similarity 0.69 | 2026 | Green |
| Safety-Constrained Agentic AI for Autism Screening: A Multimodal, Clinician-Guided Architecture Cureus discovered: cites 1 anchor(s) (A Practical Guide for Designing, Developing, and D); similarity 0.68 | 2026 | Green |
| Towards agentic smart design: An industrial large model-driven human-in-the-loop agentic workflow for geometric modelling Applied Soft Computing discovered: cites 1 anchor(s) (AFlow: Automating Agentic Workflow Generation); similarity 0.67 | 2025 | Red |