KRAGEN: a knowledge graph-enhanced RAG framework for biomedical problem solving using large language models
Nicholas Matsumoto, Jay Moran, Hyun‐Jun Choi, Miguel E Hernandez, Mythreye Venkatesan, Paul P. Wang, Jason H. Moore
Why it has this license class
AChecked 1 Oct 2026. Open license (CC-BY, CC-BY-SA, CC0, public domain): full text indexed and used in synthesis.
| Source | License | Open-access status | Read as |
|---|---|---|---|
| openalex | cc-by | gold | Green |
| crossref | https://creativecommons.org/licenses/by/4.0/ | — | Green |
| unpaywall | cc-by | gold | Green |
| europepmc | cc by | — | Green |
Abstract
BMOTIVATION: Answering and solving complex problems using a large language model (LLM) given a certain domain such as biomedicine is a challenging task that requires both factual consistency and logic, and LLMs often suffer from some major limitations, such as hallucinating false or irrelevant information, or being influenced by noisy data. These issues can compromise the trustworthiness, accuracy, and compliance of LLM-generated text and insights. RESULTS: Knowledge Retrieval Augmented Generation ENgine (KRAGEN) is a new tool that combines knowledge graphs, Retrieval Augmented Generation (RAG), and advanced prompting techniques to solve complex problems with natural language. KRAGEN converts knowledge graphs into a vector database and uses RAG to retrieve relevant facts from it. KRAGEN uses advanced prompting techniques: namely graph-of-thoughts (GoT), to dynamically break down a complex problem into smaller subproblems, and proceeds to solve each subproblem by using the relevant knowledge through the RAG framework, which limits the hallucinations, and finally, consolidates the subproblems and provides a solution. KRAGEN's graph visualization allows the user to interact with and evaluate the quality of the solution's GoT structure and logic. AVAILABILITY AND IMPLEMENTATION: KRAGEN is deployed by running its custom Docker containers. KRAGEN is available as open-source from GitHub at: https://github.com/EpistasisLab/KRAGEN.
Claims built on this paper
D- Hallucination and factual inaccuracy are the shared motivation for grounding LLMs across this body of work: Akari et al. 2023 blame errors on sole reliance on parametric knowledge, Matsumoto et al. 2024 cite hallucinated or irrelevant content and noisy data, and Sušnjak et al. 2025 explicitly design hallucination-mitigation solutions. Trace →
- Both Akari et al. 2023 and Matsumoto et al. 2024 warn that naive retrieval creates its own reliability problems—unhelpful output from indiscriminately injected passages, and difficulty selecting appropriate knowledge from large, noisy sources—motivating their respective self-reflection and graph-of-thoughts mechanisms. Trace →
- Hybrid retrieval that combines vector or semantic search with keyword matching or knowledge-graph structure appears independently in fire investigation (Choi & Cho 2026), smart manufacturing (Wan et al. 2025), and biomedicine (Matsumoto et al. 2024), suggesting a convergent design pattern for making domain RAG systems reliable. Trace →
- Traceability of generated text back to its sources emerges as a shared reliability mechanism: Sušnjak et al. 2025 propose tracking LLM responses to their information sources, and Matsumoto et al. 2024 envision medical deployments of KRAGEN providing personalized, evidence-based solutions with full transparency of reasoning and knowledge. Trace →