The role of large language models in the peer-review process: opportunities and challenges for medical journal reviewers and editors
Jisoo Lee, Jieun Lee, Jeong‐Ju Yoo
Why it has this license class
AChecked 30 Sept 2026. Open license (CC-BY, CC-BY-SA, CC0, public domain): full text indexed and used in synthesis.
| Source | License | Open-access status | Read as |
|---|---|---|---|
| openalex | cc-by | diamond | Green |
| crossref | http://creativecommons.org/licenses/by/4.0/ | — | Green |
| unpaywall | cc-by | gold | Green |
| europepmc | cc by | — | Green |
Abstract
BThe peer review process ensures the integrity of scientific research. This is particularly important in the medical field, where research findings directly impact patient care. However, the rapid growth of publications has strained reviewers, causing delays and potential declines in quality. Generative artificial intelligence, especially large language models (LLMs) such as ChatGPT, may assist researchers with efficient, high-quality reviews. This review explores the integration of LLMs into peer review, highlighting their strengths in linguistic tasks and challenges in assessing scientific validity, particularly in clinical medicine. Key points for integration include initial screening, reviewer matching, feedback support, and language review. However, implementing LLMs for these purposes will necessitate addressing biases, privacy concerns, and data confidentiality. We recommend using LLMs as complementary tools under clear guidelines to support, not replace, human expertise in maintaining rigorous peer review standards.
Claims built on this paper
D- Lee et al. (2025) note that LLM use in peer review varies by field: a corpus-level analysis found 6.5–16.9% LLM modification in AI conference reviews but no significant evidence in Nature journals. Trace →
- Kobak et al. (2025) argue their method avoids reliance on ground-truth LLM-generated text, which may not represent real scholarly LLM use, and that extending it to earlier years lets them place LLM effects in historical context against events like COVID-19. Trace →
- Estimated LLM uptake varies widely by context. Kobak et al. 2025 find differences across disciplines, countries, and journals, reaching 40% in some subcorpora. He & Bu 2026 find the highest growth in non-English-speaking countries and physical sciences. Lee et al. 2025 cite 6.5% to 16.9% LLM modification of AI-conference peer reviews but no significant evidence in Nature journals. Trace →
- Peer review is itself a site of LLM use. Resnik & Hosseini 2024 list reviewing research outputs among AI's scientific tasks. Lee et al. 2025 find LLMs strong at linguistic tasks but challenged in assessing scientific validity, particularly in clinical medicine. Trace →
- LLM adoption varies sharply across academic cultures: Lee et al. 2025 report detection estimates of 6.5–16.9% LLM-modified text in AI conference peer reviews but no significant evidence in Nature journals, mirroring the discipline, country, and journal-level variation Kobak et al. 2025 found in biomedical abstracts. Trace →
- Bechky & Davis 2024 warn that a generative AI-fueled article explosion will further strain an already overcapacity publication system, which connects to Lee et al. 2025's observation that publication growth has already strained reviewers—motivating proposals to use LLMs themselves in peer review, potentially creating a feedback loop of machine-written and machine-reviewed science. Trace →
- Lee et al. 2025 report LLM modification rates of 6.5% to 16.9% in AI conference peer reviews but no significant evidence of LLM-based modifications in Nature journals, a discipline-level variation that parallels Kobak et al. 2025's finding that LLM usage lower bounds differ across disciplines, countries, and journals. Trace →
- LLM penetration varies sharply across fields at every stage of research: Lee et al. 2025 report corpus-level detection of 6.5–16.9% LLM-modified peer reviews at AI conferences but no significant evidence of LLM-based modifications in Nature journals, mirroring Kobak et al. 2025's observation that LLM-writing prevalence differs across disciplines, countries, and journals. Trace →
- LLM use has also entered peer review, not just manuscript writing: Lee et al. 2025 report corpus-level estimates of LLM-modified AI conference reviews, and She 2026 notes that such usage spikes as deadlines approach, revealing behavioral pressures behind adoption. Trace →
- LLM use in peer review appears uneven across venues and deadline-driven: Lee et al. 2025 report detected LLM modification rates of 6.5–16.9% in AI conference reviews but no significant signal in Nature journals, while She 2026 notes that LLM-generated reviews spike as deadlines approach. Trace →