RAG reliability and evaluation
Retrieval-augmented generation and how to evaluate it: retrieval quality, citation quality, factuality. Carries the verification discipline in act 2.
Synthesised claims
AEach claim cites its evidence; a second model has checked it against that evidence.
- 01
Liu et al. 2025's systematic review and meta-analysis of 20 biomedical studies found that retrieval-augmented generation produces a statistically significant, though modest, performance improvement over baseline LLMs, with a pooled odds ratio of 1.35.
- 02
The significant average benefit of RAG quantified by Liu et al. 2025 is consistent with domain-specific evaluations: Wan et al. 2025 report 77.8% exact-match accuracy for hybrid RAG in manufacturing QA, and Gaber et al. 2025 benchmarked an LLM workflow incorporating RAG on 2,000 MIMIC-derived medical cases for triage, referral, and diagnosis support.
- 03
Hallucination and factual inaccuracy are the shared motivation for grounding LLMs across this body of work: Akari et al. 2023 blame errors on sole reliance on parametric knowledge, Matsumoto et al. 2024 cite hallucinated or irrelevant content and noisy data, and Sušnjak et al. 2025 explicitly design hallucination-mitigation solutions.
- 04
Both Akari et al. 2023 and Matsumoto et al. 2024 warn that naive retrieval creates its own reliability problems—unhelpful output from indiscriminately injected passages, and difficulty selecting appropriate knowledge from large, noisy sources—motivating their respective self-reflection and graph-of-thoughts mechanisms.
- 05
Hybrid retrieval that combines vector or semantic search with keyword matching or knowledge-graph structure appears independently in fire investigation (Choi & Cho 2026), smart manufacturing (Wan et al. 2025), and biomedicine (Matsumoto et al. 2024), suggesting a convergent design pattern for making domain RAG systems reliable.
- 06
Chen et al. 2026 argue that keyword- and rule-based clinical text extraction breaks down on negation, temporal reasoning, and cross-paragraph dependencies, and that LLMs paired with retrieval augmentation and structured output constraints enable a "generation-as-structured-output" paradigm—though they hedge that RAG only "may improve" factual errors and inconsistencies.
- 07
Sušnjak et al. 2025 show that fine-tuned open-source LLMs can automate the knowledge-discovery and synthesis phases of systematic literature reviews while maintaining high factual fidelity, validated through the replication of an existing PRISMA-conforming review.
- 08
Traceability of generated text back to its sources emerges as a shared reliability mechanism: Sušnjak et al. 2025 propose tracking LLM responses to their information sources, and Matsumoto et al. 2024 envision medical deployments of KRAGEN providing personalized, evidence-based solutions with full transparency of reasoning and knowledge.
- 09
Liu et al. 2025 identify integrating RAG systems within electronic health records as a key future direction, which speaks directly to the persistent "last-mile" bottleneck in preparing research-grade clinical datasets that Chen et al. 2026 describe.
Anchor papers
B| Paper | Year | License |
|---|---|---|
| RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation arXiv· preprint From the reviewed export of the private instance (2026-09-30): identifier only. | 2026 | Orange |
| Retrieval-Augmented Generation for Large Language Models: A Survey arXiv (Cornell University)· preprint From the reviewed export of the private instance (2026-09-30): identifier only. | 2023 | Orange |
| CareerX: A Retrieval-Augmented Generation Framework for Personalized AI-Driven Career Guidance arXiv (Cornell University)· preprint From the reviewed export of the private instance (2026-09-30): identifier only. | 2023 | Green |
| Affordance-Compiled Intelligence: Observable-Only Cognitive Impedance Matching for No-Meta LLM-Integrated Systems arXiv (Cornell University)· preprint From the reviewed export of the private instance (2026-09-30): identifier only. | 2020 | Orange |
Found through citations
C| Paper | Year | License |
|---|---|---|
| Design and Implementation of Large Language Model-Based Inference Pipeline for Fire Investigation Fire science and engineering discovered: cites 1 anchor(s) (CareerX: A Retrieval-Augmented Generation Framewor); similarity 0.71 | 2026 | Yellow |
| Operationalizing Large Language Models for Clinical Research Data Extraction: Methods, Quality Control, and Governance Journal of Medical Systems discovered: cites 1 anchor(s) (CareerX: A Retrieval-Augmented Generation Framewor); similarity 0.69 | 2026 | Green |
| Evaluating Retrieval Augmented Generation: A Comprehensive Review of Evaluation Dimensions, Question Types, and Application SN Computer Science discovered: cites 1 anchor(s) (CareerX: A Retrieval-Augmented Generation Framewor); similarity 0.69 | 2026 | Green |
| iRAT: Replanning and Controlled Retrieval for Robust LLM Reasoning Preprints.org· preprint discovered: cites 1 anchor(s) (CareerX: A Retrieval-Augmented Generation Framewor); similarity 0.66 | 2025 | Green |
| A comprehensive survey on integrating large language models with knowledge-based methods Knowledge-Based Systems discovered: cites 1 anchor(s) (Retrieval-Augmented Generation for Large Language ); similarity 0.66 | 2025 | Yellow conflict |
| Automating Research Synthesis with Domain-Specific Large Language Model Fine-Tuning ACM Transactions on Knowledge Discovery from Data discovered: cites 1 anchor(s) (CareerX: A Retrieval-Augmented Generation Framewor); similarity 0.70 | 2025 | Yellow conflict |
| Advancing Large Language Models with Enhanced Retrieval-Augmented Generation: Evidence from Biological UAV Swarm Control Drones discovered: cites 1 anchor(s) (CareerX: A Retrieval-Augmented Generation Framewor); similarity 0.68 | 2025 | Green |
| Evaluating large language model workflows in clinical decision support for triage and referral and diagnosis npj Digital Medicine discovered: cites 1 anchor(s) (Affordance-Compiled Intelligence: Observable-Only ); similarity 0.68 | 2025 | Green |
| Improving large language model applications in biomedicine with retrieval-augmented generation: a systematic review, meta-analysis, and clinical development guidelines Journal of the American Medical Informatics Association discovered: cites 1 anchor(s) (Retrieval-Augmented Generation for Large Language ); similarity 0.68 | 2025 | Yellow |
| Multi-objective math problem generation using large language model through an adaptive multi-level retrieval augmentation framework Information Fusion discovered: cites 1 anchor(s) (CareerX: A Retrieval-Augmented Generation Framewor); similarity 0.68 | 2025 | Red |
| A Survey of Conversational Search ACM Transactions on Information Systems discovered: cites 1 anchor(s) (Retrieval-Augmented Generation for Large Language ); similarity 0.67 | 2025 | Green |
| ChatCNC: Conversational machine monitoring via large language model and real-time data retrieval augmented generation Journal of Manufacturing Systems discovered: cites 1 anchor(s) (Retrieval-Augmented Generation for Large Language ); similarity 0.67 | 2025 | Red |
| Empowering LLMs by hybrid retrieval-augmented generation for domain-centric Q&A in smart manufacturing Advanced Engineering Informatics discovered: cites 1 anchor(s) (Affordance-Compiled Intelligence: Observable-Only ); similarity 0.67 | 2025 | Yellow |
| LLM Fine-Tuning: Concepts, Opportunities, and Challenges Big Data and Cognitive Computing discovered: cites 1 anchor(s) (Affordance-Compiled Intelligence: Observable-Only ); similarity 0.67 | 2025 | Green |
| From Matching to Generation: A Survey on Generative Information Retrieval ACM Transactions on Information Systems discovered: cites 1 anchor(s) (CareerX: A Retrieval-Augmented Generation Framewor); similarity 0.66 | 2025 | Red |
| A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions ACM Transactions on Information Systems discovered: cites 1 anchor(s) (CareerX: A Retrieval-Augmented Generation Framewor); similarity 0.66 | 2024 | Red |
| Advancing Retrieval-Augmented Generation with Inverted Question Matching for Enhanced QA Performance IEEE Access discovered: cites 1 anchor(s) (Retrieval-Augmented Generation for Large Language ); similarity 0.69 | 2024 | Green |
| CRP-RAG: A Retrieval-Augmented Generation Framework for Supporting Complex Logical Reasoning and Knowledge Planning Electronics discovered: cites 1 anchor(s) (CareerX: A Retrieval-Augmented Generation Framewor); similarity 0.70 | 2024 | Green |
| Business insights using RAG–LLMs: a review and case study Journal of Decision System discovered: cites 2 anchor(s) (CareerX: A Retrieval-Augmented Generation Framewor; Retrieval-Augmented Generation for Large Language ); similarity 0.68 | 2024 | Green conflict |
| Large Language Models in Medicine: The Potentials and Pitfalls Annals of Internal Medicine discovered: cites 1 anchor(s) (Affordance-Compiled Intelligence: Observable-Only ); similarity 0.67 | 2024 | Yellow conflict |
| KRAGEN: a knowledge graph-enhanced RAG framework for biomedical problem solving using large language models Bioinformatics discovered: cites 1 anchor(s) (Affordance-Compiled Intelligence: Observable-Only ); similarity 0.66 | 2024 | Green |
| Toward Robust RALMs: Revealing the Impact of Imperfect Retrieval on Retrieval-Augmented Language Models Transactions of the Association for Computational Linguistics discovered: cites 1 anchor(s) (CareerX: A Retrieval-Augmented Generation Framewor); similarity 0.67 | 2024 | Green |
| GeneGPT: augmenting large language models with domain tools for improved access to biomedical information Bioinformatics discovered: cites 1 anchor(s) (Affordance-Compiled Intelligence: Observable-Only ); similarity 0.66 | 2024 | Green |
| AI–Human Hybrids for Marketing Research: Leveraging Large Language Models (LLMs) as Collaborators Journal of Marketing discovered: cites 1 anchor(s) (Retrieval-Augmented Generation for Large Language ); similarity 0.66 | 2024 | Red |
| Materials science in the era of large language models: a perspective Digital Discovery discovered: cites 1 anchor(s) (Retrieval-Augmented Generation for Large Language ); similarity 0.65 | 2024 | Green |