Evaluating large language model workflows in clinical decision support for triage and referral and diagnosis
Farieda Gaber, Maqsood Shaik, Fabio Allega, Agnes Julia Bilecz, Felix Busch, Kelsey Goon, Vedran Franke, Altuna Akalin
Why it has this license class
AChecked 30 Sept 2026. Open license (CC-BY, CC-BY-SA, CC0, public domain): full text indexed and used in synthesis.
| Source | License | Open-access status | Read as |
|---|---|---|---|
| openalex | cc-by | gold | Green |
| crossref | https://creativecommons.org/licenses/by/4.0 | — | Green |
| unpaywall | cc-by | gold | Green |
| europepmc | cc by | — | Green |
Abstract
BAccurate medical decision-making is critical for both patients and clinicians. Patients often struggle to interpret their symptoms, determine their severity, and select the right specialist. Simultaneously, clinicians face challenges in integrating complex patient data to make timely, accurate diagnoses. Recent advances in large language models (LLMs) offer the potential to bridge this gap by supporting decision-making for both patients and healthcare providers. In this study, we benchmark multiple LLM versions and an LLM-based workflow incorporating retrieval-augmented generation (RAG) on a curated dataset of 2000 medical cases derived from the Medical Information Mart for Intensive Care database. Our findings show that these LLMs are capable of providing personalized insights into likely diagnoses, suggesting appropriate specialists, and assessing urgent care needs. These models may also support clinicians in refining diagnoses and decision-making, offering a promising approach to improving patient outcomes and streamlining healthcare delivery.
Claims built on this paper
D- The significant average benefit of RAG quantified by Liu et al. 2025 is consistent with domain-specific evaluations: Wan et al. 2025 report 77.8% exact-match accuracy for hybrid RAG in manufacturing QA, and Gaber et al. 2025 benchmarked an LLM workflow incorporating RAG on 2,000 MIMIC-derived medical cases for triage, referral, and diagnosis support. Trace →