Fine-Grained Detection of AI-Generated Writing in the Biomedical Literature
Richard She
Why it has this license class
AChecked 1 Oct 2026. Open license (CC-BY, CC-BY-SA, CC0, public domain): full text indexed and used in synthesis.
| Source | License | Open-access status | Read as |
|---|---|---|---|
| openalex | cc-by | green | Green |
| crossref | http://creativecommons.org/licenses/by/4.0/ | — | Green |
| unpaywall | cc-by | green | Green |
Abstract
BAbstract Generative AI systems are rapidly being integrated into scientific workflows, yet the specific ways in which AI-generated prose appears in published literature remain poorly characterized. Here, we use Pangram, a transformer-based detector optimized for adversarial paraphrasing, to analyze full-length biomedical research articles from 13 major journals. Papers published in 2021-2024 showed almost no detectable AI-generated text, whereas manuscripts published in 2025 exhibited a sharp increase, with 12.4% containing at least one localized passage classified as AI-written. AI usage was highly nonuniform across authors and geography: 32% of papers originating from South Korean institutions and 26% papers from Chinese institutions contained AI-generated passages, compared to 7.4% from U.S. institutions. In a focused case analysis, six labs that published fully AI-generated manuscripts also produced additional papers with extensive AI-generated segments. Journals likewise differed, with high-selectivity venues rarely containing AI-authored prose, while high-volume journals accounted for most AI-positive manuscripts. Together, these findings provide the first detailed empirical map of how and where AI-generated writing is entering the scientific literature, underscoring the need for clear norms and policies governing the use of generative AI in scientific communication.
Claims built on this paper
D- Because individual cases of AI use in papers have proven impossible to adjudicate, researchers have turned to corpus-level statistical fingerprints: She 2026 frames this necessity, and Kobak et al. 2025's method implements it by detecting emerging LLM fingerprints directly from published abstracts rather than from ground-truth datasets. Trace →
- LLM use has also entered peer review, not just manuscript writing: Lee et al. 2025 report corpus-level estimates of LLM-modified AI conference reviews, and She 2026 notes that such usage spikes as deadlines approach, revealing behavioral pressures behind adoption. Trace →
- LLM use in peer review appears uneven across venues and deadline-driven: Lee et al. 2025 report detected LLM modification rates of 6.5–16.9% in AI conference reviews but no significant signal in Nature journals, while She 2026 notes that LLM-generated reviews spike as deadlines approach. Trace →
- There is a fairness tension between detection and equity: Bietti & Bangerter 2026 note LLMs can support non-native English speakers by reducing linguistic barriers, yet the detection tools used to police LLM use are biased against non-native English writers (Liang et al., cited in She 2026). Trace →