Core bundle

Production agentic workflows and evaluation

What makes an agentic workflow fit for daily operation: evaluation beyond task success, deployment practice, and the failure modes of autonomous agents.

Synthesised claims

A

Each claim cites its evidence; a second model has checked it against that evidence.

  1. 01

    Mehta 2025 argues that current agentic AI benchmarks predominantly measure task-completion accuracy while overlooking the requirements that matter for enterprise deployment, such as cost-efficiency, reliability, and operational stability.

    Supported1 sources · 1 papersTrace →
  2. 02

    Mehta 2025 shows that single-run accuracy overstates real reliability: agent performance drops from 60% on a single run to 25% when 8-run consistency is required.

    Supported1 sources · 1 papersTrace →
  3. 03

    To close these gaps, Mehta 2025 proposes the CLEAR framework (Cost, Latency, Efficacy, Assurance, Reliability), and reports from an evaluation of six leading agents on 300 enterprise tasks that optimizing for accuracy alone yields agents 4.4–10.8x more expensive than cost-aware alternatives with comparable performance.

    Supported2 sources · 1 papersTrace →
  4. 04

    Eranga et al. 2025 characterize production agentic workflows as dynamic pipelines of multiple specialized agents using different LLMs, tool-augmented capabilities, and orchestration logic, and they offer a structured engineering lifecycle plus nine core best practices—including tool-first design over MCP and single-responsibility agents—for building them.

    Supported3 sources · 1 papersTrace →
  5. 05

    Eranga et al. 2025 present their containerized, Kubernetes-orchestrated, MCP-accessible system with Responsible-AI mechanisms as a robust template that organizations can adapt across domains such as compliance automation, media generation, analytics, and enterprise RPA.

    Supported2 sources · 1 papersTrace →
  6. 06

    Patra et al. 2026 build governance directly into their autism-screening architecture: an orchestration layer coordinates specialized agents for consent verification, bias and applicability checks, and confidence evaluation, deliberately kept as separate responsibilities so that every decision remains transparent and auditable.

    Partly supported3 sources · 1 papersTrace →
  7. 07

    Patra et al. 2026 acknowledge that their system is not yet production-validated: its abstention and governance behavior was evaluated only in a fully controlled environment, and how the system would integrate into existing clinical workflows remains unevaluated.

    Supported2 sources · 1 papersTrace →
  8. 08

    Eranga et al. 2025 and Patra et al. 2026 independently converge on separation of concerns as the foundation of auditable production agentic systems: Eranga prescribes single-responsibility agents and clean separation between workflow logic and MCP servers, while Patra assigns each architectural layer a clear responsibility so governance checks stay visible and testable.

    Supported2 sources · 2 papersTrace →
  9. 09

    The evaluation gap that Mehta 2025 quantifies in benchmarks—missing multidimensional metrics and inadequate reliability assessment—is illustrated concretely by Patra et al. 2026, whose governance and abstention behavior was validated only in a controlled environment rather than under real-world deployment conditions.

    Supported2 sources · 2 papersTrace →
  10. 10

    MCP is emerging as shared infrastructure across the production-agentic literature: Eranga et al. 2025 prescribe tool-first design over MCP with clean separation between workflow logic and MCP servers, while Yue et al. 2026's perspective on MCP-native AI scientist ecosystems cites MCP construction tooling such as Code2MCP alongside work on measuring agents in production.

    Partly supported3 sources · 2 papersTrace →

Anchor papers

B
PaperYearLicense
Agents of Chaos
arXiv (Cornell University)· preprint
From the reviewed export of the private instance (2026-09-30): identifier only.
2026Orange
Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
arXiv (Cornell University)· preprint
From the reviewed export of the private instance (2026-09-30): identifier only.
2025Green
A Practical Guide for Designing, Developing, and Deploying Production-Grade Agentic AI Workflows
arXiv (Cornell University)· preprint
From the reviewed export of the private instance (2026-09-30): identifier only.
2025Green
Orchestrating Human-AI Teams: The Manager Agent as a Unifying Research Challenge
arXiv (Cornell University)· preprint
From the reviewed export of the private instance (2026-09-30): identifier only.
2025Orange
AFlow: Automating Agentic Workflow Generation
arXiv (Cornell University)· preprint
From the reviewed export of the private instance (2026-09-30): identifier only.
2024Orange

Found through citations

C
PaperYearLicense
Building MCP-native hierarchical AI scientist ecosystems: a perspective on scaling multi-agent scientific discovery
Frontiers in Artificial Intelligence
discovered: cites 1 anchor(s) (AFlow: Automating Agentic Workflow Generation); similarity 0.69
2026Green
Safety-Constrained Agentic AI for Autism Screening: A Multimodal, Clinician-Guided Architecture
Cureus
discovered: cites 1 anchor(s) (A Practical Guide for Designing, Developing, and D); similarity 0.68
2026Green
Towards agentic smart design: An industrial large model-driven human-in-the-loop agentic workflow for geometric modelling
Applied Soft Computing
discovered: cites 1 anchor(s) (AFlow: Automating Agentic Workflow Generation); similarity 0.67
2025Red