Delving into LLM-assisted writing in biomedical publications through excess vocabulary
Dmitry Kobak, Rita González-Márquez, Emőke-Ágnes Horvát, Jan Lause
Why it has this license class
AChecked 30 Sept 2026. Open license (CC-BY, CC-BY-SA, CC0, public domain): full text indexed and used in synthesis.
| Source | License | Open-access status | Read as |
|---|---|---|---|
| openalex | cc-by | gold | Green |
| unpaywall | cc-by | gold | Green |
| europepmc | cc by | — | Green |
Abstract
BLarge language models (LLMs) like ChatGPT can generate and revise text with human-level performance. These models come with clear limitations, can produce inaccurate information, and reinforce existing biases. Yet, many scientists use them for their scholarly writing. But how widespread is such LLM usage in the academic literature? To answer this question for the field of biomedical research, we present an unbiased, large-scale approach: We study vocabulary changes in more than 15 million biomedical abstracts from 2010 to 2024 indexed by PubMed and show how the appearance of LLMs led to an abrupt increase in the frequency of certain style words. This excess word analysis suggests that at least 13.5% of 2024 abstracts were processed with LLMs. This lower bound differed across disciplines, countries, and journals, reaching 40% for some subcorpora. We show that LLMs have had an unprecedented impact on scientific writing in biomedical research, surpassing the effect of major world events such as the COVID pandemic.
Claims built on this paper
D- Kobak et al. (2025) use excess-vocabulary analysis of PubMed abstracts to estimate that at least 13.5% of 2024 biomedical abstracts were processed with LLMs, with some subcorpora reaching 40%. Trace →
- Kobak et al. (2025) argue their method avoids reliance on ground-truth LLM-generated text, which may not represent real scholarly LLM use, and that extending it to earlier years lets them place LLM effects in historical context against events like COVID-19. Trace →
- Fortenbach et al. (2025/2026) find the same kind of stylistic-word surge in ophthalmology that Kobak et al. (2025) found in biomedicine, and additionally use AI detectors to report that 25.7% of sampled research articles had outlier AI-likelihood scores by 2025. Trace →
- Bietti & Bangerter (2026) cite prevalence estimates of AI-generated text in biomedical papers, consistent with the excess-vocabulary work, and describe LLM use as swift and largely unregulated, shifting debate toward disclosure norms and accountability. Trace →
- Two corpus-level studies measure LLM use through surges in stylistic words. Kobak et al. 2025 estimate that at least 13.5% of 2024 PubMed abstracts were LLM-processed, and Fortenbach et al. 2026 find at least a 2-fold usage increase in 20% of ophthalmology abstracts after ChatGPT's release. Trace →
- Estimated LLM uptake varies widely by context. Kobak et al. 2025 find differences across disciplines, countries, and journals, reaching 40% in some subcorpora. He & Bu 2026 find the highest growth in non-English-speaking countries and physical sciences. Lee et al. 2025 cite 6.5% to 16.9% LLM modification of AI-conference peer reviews but no significant evidence in Nature journals. Trace →
- Language support is a recurring rationale for LLM use and may bear on where adoption grows fastest. Kobak et al. 2025 and Bietti & Bangerter 2026 both note help for writing in English. He & Bu 2026 observe the highest growth in non-English-speaking countries. Trace →
- Detecting LLM text is methodologically fraught. Bietti & Bangerter 2026 note that detection systems face significant limitations. Kobak et al. 2025 argue that prior detection studies relied on potentially biased ground-truth corpora, which their direct excess-vocabulary approach avoids. Trace →
- Walters & Wilder 2023 report that the fabricated share of ChatGPT-generated citations typically falls in the 47–69% range. They argue that generative AI output does not merit the unchecked trust routinely given to other research software. Trace →
- Kobak et al. 2025 measured LLM use in biomedical writing at an unprecedented scale, analyzing over 15 million PubMed abstracts from 2010–2024 and estimating that at least 13.5% of 2024 abstracts were processed with LLMs, with the lower bound reaching 40% in some subcorpora. Trace →
- Kobak et al. 2025 argue their method is conceptually stronger than prior detection work because it detects LLM fingerprints directly from published abstracts rather than relying on potentially biased ground-truth datasets, allowing them to conclude that LLMs' impact on scientific writing surpasses even that of the COVID pandemic. Trace →
- Independent detection approaches in different medical fields converge on the same pattern: the corpus-wide excess-vocabulary method of Kobak et al. 2025 and the journal-level AI-detection screening of Fortenbach et al. 2026 both document a sharp post-ChatGPT rise in LLM-generated text, with Fortenbach et al. finding over a quarter of sampled ophthalmology research articles showing outlier AI-likelihood scores by 2025. Trace →
- LLM adoption varies sharply across academic cultures: Lee et al. 2025 report detection estimates of 6.5–16.9% LLM-modified text in AI conference peer reviews but no significant evidence in Nature journals, mirroring the discipline, country, and journal-level variation Kobak et al. 2025 found in biomedical abstracts. Trace →
- Kobak et al. 2025 estimate that at least 13.5% of 2024 biomedical abstracts were processed with LLMs, with the lower bound reaching 40% for some subcorpora; they describe the resulting impact on scientific writing as unprecedented, surpassing the effect of major world events such as the COVID pandemic. Trace →
- Kobak et al. 2025 and Fortenbach et al. 2026 both detect LLM-assisted writing through style-word frequencies: Kobak et al. studied vocabulary changes in more than 15 million PubMed abstracts, while Fortenbach et al. evaluated abstract text from 27,142 ophthalmology research articles for changes in word-frequency usage with a focus on stylistic words. Trace →
- Bietti & Bangerter 2026 report that a recent study estimated more than one in 10 biomedical papers published in 2024 contained AI-generated text, often without disclosure, consistent with Kobak et al. 2025's lower bound of at least 13.5% of 2024 abstracts. They argue that using such tools can risk decoupling writing from thinking. Trace →
- Lee et al. 2025 report LLM modification rates of 6.5% to 16.9% in AI conference peer reviews but no significant evidence of LLM-based modifications in Nature journals, a discipline-level variation that parallels Kobak et al. 2025's finding that LLM usage lower bounds differ across disciplines, countries, and journals. Trace →
- The language-equity rationale for LLM-assisted writing recurs across papers: Kobak et al. 2025 note that LLMs can help translate to English, Bietti & Bangerter 2026 note that LLMs can support non-native English speakers by reducing linguistic barriers, and He & Bu 2026 find that non-English-speaking countries exhibit the highest growth rates in AI writing tool use. Trace →
- Kobak et al. 2025's excess-vocabulary analysis of more than 15 million PubMed-indexed biomedical abstracts (2010–2024) suggests that at least 13.5% of 2024 abstracts were processed with LLMs, a lower bound reaching 40% for some subcorpora; the authors judge the resulting stylistic shift unprecedented, surpassing even the effect of the COVID pandemic. Trace →
- LLM penetration varies sharply across fields at every stage of research: Lee et al. 2025 report corpus-level detection of 6.5–16.9% LLM-modified peer reviews at AI conferences but no significant evidence of LLM-based modifications in Nature journals, mirroring Kobak et al. 2025's observation that LLM-writing prevalence differs across disciplines, countries, and journals. Trace →
- Fortenbach et al. 2026 transport Kobak et al. 2025's excess-vocabulary method into a single clinical specialty, screening ophthalmology abstracts for changes in word-frequency usage focused on stylistic words previously associated with LLM-generated text — the same abrupt style-word increases Kobak et al. used to date LLM adoption across 15 million biomedical abstracts. Trace →
- Bietti & Bangerter 2026 anchor their warning that LLMs risk decoupling writing from thinking in prevalence estimates — more than one in ten biomedical papers published in 2024 and over 20% in some fields — consistent with the lower bounds Kobak et al. 2025 derived from excess vocabulary, showing how measurement studies now feed the normative debate. Trace →
- Bechky & Davis 2024, writing from a sociology-of-science standpoint, illustrate undisclosed LLM writing with an article that opened "Certainly, here is a possible introduction to your topic" and point readers to Kobak and colleagues' excess-vocabulary research for a systematic assessment in the medical literature, showing the measurement approach had already crossed disciplinary boundaries. Trace →
- Fortenbach et al. 2026 find that LLM-generated text is now common in ophthalmology: by 2025, 25.7% of sampled research articles and 21.6% of commentary articles carried AI-likelihood scores more than two standard deviations above baseline, and none of these outlier publications disclosed AI use. Trace →
- Independent detection methods converge in showing that LLM-generated text has rapidly entered the biomedical literature: Kobak et al. 2025's excess-vocabulary analysis of over 15 million PubMed abstracts and Fortenbach et al. 2026's detector-based screening of ophthalmology journals both document steep increases after ChatGPT's release. Trace →
- Because individual cases of AI use in papers have proven impossible to adjudicate, researchers have turned to corpus-level statistical fingerprints: She 2026 frames this necessity, and Kobak et al. 2025's method implements it by detecting emerging LLM fingerprints directly from published abstracts rather than from ground-truth datasets. Trace →
- Kobak et al. 2025 analyzed more than 15 million biomedical abstracts from 2010 to 2024 and estimated that at least 13.5% of 2024 abstracts were processed with LLMs, with the lower bound reaching 40% in some subcorpora. Trace →
- Field-specific screening corroborates the biomedical-wide estimate: Fortenbach et al. 2026 found that by 2025 about a quarter of sampled ophthalmology research articles carried AI-likelihood scores more than two standard deviations above baseline, consistent with the lower bound reported by Kobak et al. 2025. Trace →