Research integrity and LLMs
What LLMs do to the evidence base: illusions of understanding, fabricated references, and measurable LLM text in papers and peer review. Carries act 2.
Synthesised claims
AEach claim cites its evidence; a second model has checked it against that evidence.
- 01
Kobak et al. 2025 analyzed more than 15 million biomedical abstracts from 2010 to 2024 and estimated that at least 13.5% of 2024 abstracts were processed with LLMs, with the lower bound reaching 40% in some subcorpora.
- 02
Field-specific screening corroborates the biomedical-wide estimate: Fortenbach et al. 2026 found that by 2025 about a quarter of sampled ophthalmology research articles carried AI-likelihood scores more than two standard deviations above baseline, consistent with the lower bound reported by Kobak et al. 2025.
- 03
Disclosure is almost entirely absent even where LLM use is detected: He & Bu 2026 found only about 0.1% of post-2023 papers disclosed AI use despite most journals having policies, and Fortenbach et al. 2026 found no disclosure among any publications with outlier scores, indicating current policies fail to promote transparency.
- 04
Walters & Wilder 2023 document that ChatGPT fabricates a large share of the bibliographic citations it generates, with systematic studies typically finding fabrication rates in the 47–69% range.
- 05
LLM use in peer review appears uneven across venues and deadline-driven: Lee et al. 2025 report detected LLM modification rates of 6.5–16.9% in AI conference reviews but no significant signal in Nature journals, while She 2026 notes that LLM-generated reviews spike as deadlines approach.
- 06
Two commentaries converge on an epistemic risk of LLM adoption in science: Messeri & Crockett 2024 warn of a phase in which we produce more but understand less, and Bietti & Bangerter 2026 argue that outsourcing writing to LLMs risks decoupling writing from thinking.
- 07
There is a fairness tension between detection and equity: Bietti & Bangerter 2026 note LLMs can support non-native English speakers by reducing linguistic barriers, yet the detection tools used to police LLM use are biased against non-native English writers (Liang et al., cited in She 2026).
- 08
Undisclosed LLM use has slipped past peer review in visible ways: Bechky & Davis 2024 recount an article opening with a telltale ChatGPT phrase, and Izquierdo-Condoy et al. 2026 document stock LLM phrases appearing in papers that passed peer review.
- 09
Both Binz et al. 2025 and Walters & Wilder 2023 caution that verification costs may erode LLM efficiency gains: the trust routinely placed in ordinary software is inappropriate for generative AI, and time saved in text generation may be offset by the time required to verify the output.
- 10
Arzilli et al. 2025 warn that increased reliance on AI-driven authoring tools could exacerbate an infodemic of unreliable or misleading information, a risk amplified by the prevailing publish-or-perish culture.
- 11
Both Binz et al. 2025 and Walters & Wilder 2023 caution that verification costs may erode LLM efficiency gains: the trust routinely placed in ordinary software is inappropriate for generative AI, and time saved in text generation may be offset by the time required to verify the output.
- 12
Independent detection methods converge in showing that LLM-generated text has rapidly entered the biomedical literature: Kobak et al. 2025's excess-vocabulary analysis of over 15 million PubMed abstracts and Fortenbach et al. 2026's detector-based screening of ophthalmology journals both document steep increases after ChatGPT's release.
- 13
Although AI-assisted writing is now common, disclosure of it is nearly nonexistent: He & Bu 2026 found almost no explicit disclosures in 75,000 recent papers, and Fortenbach et al. 2026 found none at all among ophthalmology publications flagged as likely AI-written.
- 14
Journal AI policies have so far been ineffective at changing author behavior: He & Bu 2026 show that despite 70% of journals adopting AI policies, mostly requiring disclosure, AI writing tool use grew just as much in journals with policies as in those without.
- 15
Walters & Wilder 2023 document that ChatGPT systematically fabricates bibliographic citations, with the systematic studies they review typically finding fabrication rates between 47% and 69%.
- 16
Bietti & Bangerter 2026 and Messeri & Crockett 2024 articulate a shared epistemic worry: because scientific writing is itself a form of thinking, outsourcing it to LLMs risks a mode of science that produces more text while understanding less.
- 17
Because individual cases of AI use in papers have proven impossible to adjudicate, researchers have turned to corpus-level statistical fingerprints: She 2026 frames this necessity, and Kobak et al. 2025's method implements it by detecting emerging LLM fingerprints directly from published abstracts rather than from ground-truth datasets.
- 18
LLM use has also entered peer review, not just manuscript writing: Lee et al. 2025 report corpus-level estimates of LLM-modified AI conference reviews, and She 2026 notes that such usage spikes as deadlines approach, revealing behavioral pressures behind adoption.
- 19
The democratizing promise of LLMs for non-native English writers, acknowledged even by critics such as Bietti & Bangerter 2026, is consistent with He & Bu 2026's large-scale finding that non-English-speaking countries exhibit the highest growth rates in AI-assisted writing.
- 20
Arzilli et al. 2025 and Bechky & Davis 2024 both warn that generative AI will inflate publication volume and further strain an already overcapacity publishing system, compounding publish-or-perish pressures and the risk of an unreliable-information infodemic.
- 21
Binz et al. 2025 and Walters & Wilder 2023 converge on the point that generative AI cannot be trusted the way conventional research software is: the verification burden falls on authors and may offset whatever time text generation saves.
- 22
Kobak et al. 2025's excess-vocabulary analysis of more than 15 million PubMed-indexed biomedical abstracts (2010–2024) suggests that at least 13.5% of 2024 abstracts were processed with LLMs, a lower bound reaching 40% for some subcorpora; the authors judge the resulting stylistic shift unprecedented, surpassing even the effect of the COVID pandemic.
- 23
Walters & Wilder 2023 show that fabricated references are a systematic failure mode of ChatGPT-assisted writing: across studies the share of fabricated citations typically falls in the 47–69% range, and one radiology evaluation found 64% of 343 citations could not be found in PubMed or on the open web.
- 24
He & Bu 2026 conclude that journal AI policies have largely failed: analyzing 5,114 journals and over 5.2 million papers, they find 70% of journals adopted policies (primarily requiring disclosure), yet AI writing tool use rose with no significant difference between journals with and without policies, and only about 0.1% of papers published since 2023 disclosed AI use.
- 25
Fortenbach et al. 2026 find that LLM-generated text is now common in ophthalmology: by 2025, 25.7% of sampled research articles and 21.6% of commentary articles carried AI-likelihood scores more than two standard deviations above baseline, and none of these outlier publications disclosed AI use.
- 26
LLM penetration varies sharply across fields at every stage of research: Lee et al. 2025 report corpus-level detection of 6.5–16.9% LLM-modified peer reviews at AI conferences but no significant evidence of LLM-based modifications in Nature journals, mirroring Kobak et al. 2025's observation that LLM-writing prevalence differs across disciplines, countries, and journals.
- 27
Fortenbach et al. 2026 transport Kobak et al. 2025's excess-vocabulary method into a single clinical specialty, screening ophthalmology abstracts for changes in word-frequency usage focused on stylistic words previously associated with LLM-generated text — the same abrupt style-word increases Kobak et al. used to date LLM adoption across 15 million biomedical abstracts.
- 28
Independent corpora converge on an AI transparency gap: He & Bu 2026 find that among 75,000 papers published since 2023 only 76 (~0.1%) disclosed AI use, while Fortenbach et al. 2026 report that none of the ophthalmology publications with outlier AI-likelihood scores disclosed their AI use.
- 29
Bietti & Bangerter 2026 anchor their warning that LLMs risk decoupling writing from thinking in prevalence estimates — more than one in ten biomedical papers published in 2024 and over 20% in some fields — consistent with the lower bounds Kobak et al. 2025 derived from excess vocabulary, showing how measurement studies now feed the normative debate.
- 30
Two perspective pieces converge on the epistemic stakes of AI-assisted science: Messeri & Crockett 2024 warn that proliferating AI tools risk a phase of enquiry in which we produce more but understand less, while Bietti & Bangerter 2026 argue that outsourcing writing to LLMs risks decoupling writing from thinking.
- 31
Bechky & Davis 2024, writing from a sociology-of-science standpoint, illustrate undisclosed LLM writing with an article that opened "Certainly, here is a possible introduction to your topic" and point readers to Kobak and colleagues' excess-vocabulary research for a systematic assessment in the medical literature, showing the measurement approach had already crossed disciplinary boundaries.
- 32
Fortenbach et al. 2026 find that LLM-generated text is now common in ophthalmology: by 2025, 25.7% of sampled research articles and 21.6% of commentary articles carried AI-likelihood scores more than two standard deviations above baseline, and none of these outlier publications disclosed AI use.
- 33
Kobak et al. 2025 estimate that at least 13.5% of 2024 biomedical abstracts were processed with LLMs, with the lower bound reaching 40% for some subcorpora; they describe the resulting impact on scientific writing as unprecedented, surpassing the effect of major world events such as the COVID pandemic.
- 34
Walters & Wilder 2023 report that the proportion of fabricated citations in ChatGPT-generated content typically falls in the 47–69% range, citing a radiology study in which 64% of 343 citations were fabricated, i.e., could not be found in PubMed or on the open web.
- 35
He & Bu 2026 find that despite 70% of journals adopting AI policies, researchers' use of AI writing tools has increased dramatically across disciplines, with no significant difference between journals with or without policies.
- 36
Fortenbach et al. 2026 found that by 2025, 25.7% of sampled ophthalmology research articles and 21.6% of commentary articles contained AI-likelihood scores of more than 2 standard deviations above the baseline. Among publications with outlier scores, 22.3% of sentences in research articles and 90% of sentences in commentary articles were likely written by AI.
- 37
Kobak et al. 2025 and Fortenbach et al. 2026 both detect LLM-assisted writing through style-word frequencies: Kobak et al. studied vocabulary changes in more than 15 million PubMed abstracts, while Fortenbach et al. evaluated abstract text from 27,142 ophthalmology research articles for changes in word-frequency usage with a focus on stylistic words.
- 38
Non-disclosure emerges across studies using different methods: He & Bu 2026 found that of 75,000 papers published since 2023, only 76 (~0.1%) explicitly disclosed AI use, while Fortenbach et al. 2026 found that AI use was not disclosed among any ophthalmology publications with outlier AI-likelihood scores.
- 39
Bietti & Bangerter 2026 report that a recent study estimated more than one in 10 biomedical papers published in 2024 contained AI-generated text, often without disclosure, consistent with Kobak et al. 2025's lower bound of at least 13.5% of 2024 abstracts. They argue that using such tools can risk decoupling writing from thinking.
- 40
Lee et al. 2025 report LLM modification rates of 6.5% to 16.9% in AI conference peer reviews but no significant evidence of LLM-based modifications in Nature journals, a discipline-level variation that parallels Kobak et al. 2025's finding that LLM usage lower bounds differ across disciplines, countries, and journals.
- 41
Messeri & Crockett 2024 warn that the proliferation of AI tools in science risks a phase of scientific enquiry in which we produce more but understand less, anticipating Bietti & Bangerter 2026's contention that scientific writing is not merely communication but a form of thinking.
- 42
The language-equity rationale for LLM-assisted writing recurs across papers: Kobak et al. 2025 note that LLMs can help translate to English, Bietti & Bangerter 2026 note that LLMs can support non-native English speakers by reducing linguistic barriers, and He & Bu 2026 find that non-English-speaking countries exhibit the highest growth rates in AI writing tool use.
- 43
Walters & Wilder 2023 report that the proportion of fabricated citations in ChatGPT-generated content typically falls in the 47–69% range, citing a radiology study in which 64% of 343 citations were fabricated, i.e., could not be found in PubMed or on the open web.
- 44
Kobak et al. 2025 measured LLM use in biomedical writing at an unprecedented scale, analyzing over 15 million PubMed abstracts from 2010–2024 and estimating that at least 13.5% of 2024 abstracts were processed with LLMs, with the lower bound reaching 40% in some subcorpora.
- 45
Kobak et al. 2025 argue their method is conceptually stronger than prior detection work because it detects LLM fingerprints directly from published abstracts rather than relying on potentially biased ground-truth datasets, allowing them to conclude that LLMs' impact on scientific writing surpasses even that of the COVID pandemic.
- 46
Independent detection approaches in different medical fields converge on the same pattern: the corpus-wide excess-vocabulary method of Kobak et al. 2025 and the journal-level AI-detection screening of Fortenbach et al. 2026 both document a sharp post-ChatGPT rise in LLM-generated text, with Fortenbach et al. finding over a quarter of sampled ophthalmology research articles showing outlier AI-likelihood scores by 2025.
- 47
Two large studies independently document a near-total transparency gap: He & Bu 2026 found only about 0.1% of 75,000 post-2023 papers explicitly disclosed AI use, and Fortenbach et al. 2026 found no disclosure in any ophthalmology publication flagged with outlier AI scores.
- 48
He & Bu 2026 show that journal AI policies have so far been ineffective: although 70% of 5,114 analyzed journals adopted AI policies, mostly requiring disclosure, AI writing-tool use grew dramatically with no significant difference between journals with and without such policies.
- 49
Walters & Wilder 2023 document that fabricated citations are a systematic failure mode of ChatGPT, with fabrication rates typically between 47% and 69% across studies, and argue that the trust researchers reasonably place in statistical software is not warranted for generative AI because its tasks are fundamentally different.
- 50
Messeri & Crockett 2024 and Bietti & Bangerter 2026 raise complementary epistemic alarms about AI in science: the former warn of a phase of enquiry in which we produce more but understand less, while the latter argue that outsourcing writing to LLMs risks decoupling writing from thinking, since scientific writing is itself a form of thinking.
Anchor papers
B| Paper | Year | License |
|---|---|---|
| Quantifying large language model usage in scientific papers Nature Human Behaviour A measurable rise in LLM use across fields. | 2025 | Red |
| Delving into LLM-assisted writing in biomedical publications through excess vocabulary Science Advances At least 13.5% of 2024 PubMed abstracts were processed with LLMs. | 2025 | Green |
| A Critical Analysis of Generative AI: Challenges, Opportunities, and Future Research Directions Archives of Computational Methods in Engineering Added by Fredrik in Zotero, 2026-09-30: critical analysis of generative AI. | 2025 | Green |
| RETRACTED: The three-dimensional porous mesh structure of Cu-based metal-organic-framework - Aramid cellulose separator enhances the electrochemical performance of lithium metal anode batteries Surfaces and Interfaces Retracted Retracted 2024 (Retraction Watch 53671): the introduction opened with "Certainly, here is". Publisher terms only: red and retracted. | 2024 | Red |
| Artificial intelligence and illusions of understanding in scientific research Nature Illusions of understanding. The load-bearing text for act 2. | 2024 | Red |
| Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews arXiv (Cornell University)· preprint LLM-generated text in peer review. | 2024 | Orange |
| RETRACTED: Cellular functions of spermatogonial stem cells in relation to JAK/STAT signaling pathway Frontiers in Cell and Developmental Biology Retracted Retracted 2024 (Retraction Watch 51983): AI-generated figures. CC-BY yet retracted: the license says it may be stored, the retraction keeps it out of synthesis. | 2024 | Green |
| Fabrication and errors in the bibliographic citations generated by ChatGPT Scientific Reports 55% fabricated references with GPT-3.5, 18% with GPT-4. | 2023 | Green |
| AI and science: what 1,600 researchers think Nature Researchers' own concerns. | 2023 | Red |
Found through citations
C| Paper | Year | License |
|---|---|---|
| Will the widespread use of large language models in scientific writing undermine scientists’ critical thinking? PLoS Biology discovered: cites 2 anchor(s) (Delving into LLM-assisted writing in biomedical pu; Quantifying large language model usage in scientif); similarity 0.75 | 2026 | Green |
| Artificial Intelligence for Academic Text Generation in Analytical Chemistry: Current Risks, Indicators, and Perspectives toward Greener and More Sustainable Approaches Analytical Chemistry discovered: cites 2 anchor(s) (Delving into LLM-assisted writing in biomedical pu; Monitoring AI-Modified Content at Scale: A Case St); similarity 0.79 | 2026 | Red |
| Academic journals’ AI policies fail to curb the surge in AI-assisted academic writing Proceedings of the National Academy of Sciences discovered: cites 3 anchor(s) (Delving into LLM-assisted writing in biomedical pu; Quantifying large language model usage in scientif; AI and science: what 1,600 researchers think); similarity 0.73 | 2026 | Yellow |
| Large Language Model Authorship in Ophthalmic Publications Ophthalmology discovered: cites 2 anchor(s) (Delving into LLM-assisted writing in biomedical pu; Quantifying large language model usage in scientif); similarity 0.75 | 2026 | Yellow |
| Fine-Grained Detection of AI-Generated Writing in the Biomedical Literature bioRxiv (Cold Spring Harbor Laboratory)· preprint discovered: cites 2 anchor(s) (Monitoring AI-Modified Content at Scale: A Case St; Delving into LLM-assisted writing in biomedical pu); similarity 0.73 | 2026 | Green |
| Artificial Intelligence in Medical Education: Transformative Potential, Current Applications, and Future Implications JMIR Medical Education discovered: cites 1 anchor(s) (Fabrication and errors in the bibliographic citati); similarity 0.77 | 2026 | Green |
| The Epistemic Downside of Using LLM-Based Generative AI in Academic Writing Publications discovered: cites 2 anchor(s) (Delving into LLM-assisted writing in biomedical pu; Quantifying large language model usage in scientif); similarity 0.75 | 2025 | Green |
| A Primer for Evaluating Large Language Models in Social-Science Research Advances in Methods and Practices in Psychological Science discovered: cites 2 anchor(s) (Artificial intelligence and illusions of understan; Delving into LLM-assisted writing in biomedical pu); similarity 0.75 | 2025 | Yellow |
| The role of large language models in the peer-review process: opportunities and challenges for medical journal reviewers and editors Journal of Educational Evaluation for Health Professions discovered: cites 2 anchor(s) (Monitoring AI-Modified Content at Scale: A Case St; AI and science: what 1,600 researchers think); similarity 0.77 | 2025 | Green |
| Generative Artificial Intelligence (GenAI) in the research process – A survey of researchers’ practices and perceptions Technology in Society discovered: cites 2 anchor(s) (Artificial intelligence and illusions of understan; AI and science: what 1,600 researchers think); similarity 0.75 | 2025 | Green |
| How should the advancement of large language models affect the practice of science? Proceedings of the National Academy of Sciences discovered: cites 2 anchor(s) (Fabrication and errors in the bibliographic citati; Artificial intelligence and illusions of understan); similarity 0.73 | 2025 | Green |
| Generative Artificial Intelligence: Implications for Biomedical and Health Professions Education Annual Review of Biomedical Data Science discovered: cites 1 anchor(s) (Fabrication and errors in the bibliographic citati); similarity 0.77 | 2025 | Green |
| A surge of AI-driven publications: the impact on health professionals and potential mitigating solutions Frontiers in Public Health discovered: cites 1 anchor(s) (RETRACTED: Cellular functions of spermatogonial st); similarity 0.76 | 2025 | Green |
| The role of generative AI in academic and scientific authorship: an autopoietic perspective AI & Society discovered: cites 2 anchor(s) (Monitoring AI-Modified Content at Scale: A Case St; RETRACTED: Cellular functions of spermatogonial st); similarity 0.77 | 2025 | Green |
| The Use of Large Language Models and Their Association With Enhanced Impact in Biomedical Research and Beyond MedComm – Future Medicine discovered: cites 1 anchor(s) (Monitoring AI-Modified Content at Scale: A Case St); similarity 0.76 | 2025 | Green |
| “AI et al.” The perils of overreliance on Artificial Intelligence by authors in scientific research Clinical eHealth discovered: cites 2 anchor(s) (RETRACTED: Cellular functions of spermatogonial st; RETRACTED: The three-dimensional porous mesh struc); similarity 0.77 | 2024 | Yellow |
| ChatGPT in Teaching and Learning: A Systematic Review Education Sciences discovered: cites 1 anchor(s) (Fabrication and errors in the bibliographic citati); similarity 0.76 | 2024 | Green |
| Evaluating Literature Reviews Conducted by Humans Versus ChatGPT: Comparative Study JMIR AI discovered: cites 1 anchor(s) (AI and science: what 1,600 researchers think); similarity 0.77 | 2024 | Green |
| The transformative impact of large language models on medical writing and publishing: current applications, challenges and future directions Korean Journal of Physiology and Pharmacology discovered: cites 1 anchor(s) (Monitoring AI-Modified Content at Scale: A Case St); similarity 0.77 | 2024 | Yellow |
| The ethics of using artificial intelligence in scientific research: new guidance needed for a new tool AI and Ethics discovered: cites 2 anchor(s) (Artificial intelligence and illusions of understan; Fabrication and errors in the bibliographic citati); similarity 0.75 | 2024 | Green |
| ChatGPT and generative AI are revolutionizing the scientific community: A Janus‐faced conundrum iMeta discovered: cites 1 anchor(s) (AI and science: what 1,600 researchers think); similarity 0.77 | 2024 | Green |
| Resisting the Algorithmic Management of Science: Craft and Community After Generative AI Administrative Science Quarterly discovered: cites 3 anchor(s) (Artificial intelligence and illusions of understan; Delving into LLM-assisted writing in biomedical pu; Monitoring AI-Modified Content at Scale: A Case St); similarity 0.69 | 2024 | Green conflict |
| Artificial Intelligence: A Challenge to Scientific Communication Klinische Monatsblätter für Augenheilkunde discovered: cites 1 anchor(s) (RETRACTED: Cellular functions of spermatogonial st); similarity 0.79 | 2024 | Red |
| Obvious artificial intelligence ‐generated anomalies in published journal articles: A call for enhanced editorial diligence Learned Publishing discovered: cites 2 anchor(s) (RETRACTED: Cellular functions of spermatogonial st; RETRACTED: The three-dimensional porous mesh struc); similarity 0.74 | 2024 | Green |
| Implementation and Evaluation of a ChatGPT-Assisted Special Topics Writing Assignment in Biochemistry Journal of Chemical Education discovered: cites 2 anchor(s) (Fabrication and errors in the bibliographic citati; AI and science: what 1,600 researchers think); similarity 0.73 | 2024 | Red |