Core bundle

Research integrity and LLMs

What LLMs do to the evidence base: illusions of understanding, fabricated references, and measurable LLM text in papers and peer review. Carries act 2.

Synthesised claims

A

Each claim cites its evidence; a second model has checked it against that evidence.

  1. 01

    Kobak et al. 2025 analyzed more than 15 million biomedical abstracts from 2010 to 2024 and estimated that at least 13.5% of 2024 abstracts were processed with LLMs, with the lower bound reaching 40% in some subcorpora.

    Supported1 sources · 1 papersTrace →
  2. 02

    Field-specific screening corroborates the biomedical-wide estimate: Fortenbach et al. 2026 found that by 2025 about a quarter of sampled ophthalmology research articles carried AI-likelihood scores more than two standard deviations above baseline, consistent with the lower bound reported by Kobak et al. 2025.

    Partly supported2 sources · 2 papersTrace →
  3. 03

    Disclosure is almost entirely absent even where LLM use is detected: He & Bu 2026 found only about 0.1% of post-2023 papers disclosed AI use despite most journals having policies, and Fortenbach et al. 2026 found no disclosure among any publications with outlier scores, indicating current policies fail to promote transparency.

    Supported3 sources · 2 papersTrace →
  4. 04

    Walters & Wilder 2023 document that ChatGPT fabricates a large share of the bibliographic citations it generates, with systematic studies typically finding fabrication rates in the 47–69% range.

    Supported1 sources · 1 papersTrace →
  5. 05

    LLM use in peer review appears uneven across venues and deadline-driven: Lee et al. 2025 report detected LLM modification rates of 6.5–16.9% in AI conference reviews but no significant signal in Nature journals, while She 2026 notes that LLM-generated reviews spike as deadlines approach.

    Supported3 sources · 2 papersTrace →
  6. 06

    Two commentaries converge on an epistemic risk of LLM adoption in science: Messeri & Crockett 2024 warn of a phase in which we produce more but understand less, and Bietti & Bangerter 2026 argue that outsourcing writing to LLMs risks decoupling writing from thinking.

    Supported2 sources · 2 papersTrace →
  7. 07

    There is a fairness tension between detection and equity: Bietti & Bangerter 2026 note LLMs can support non-native English speakers by reducing linguistic barriers, yet the detection tools used to police LLM use are biased against non-native English writers (Liang et al., cited in She 2026).

    Supported2 sources · 2 papersTrace →
  8. 08

    Undisclosed LLM use has slipped past peer review in visible ways: Bechky & Davis 2024 recount an article opening with a telltale ChatGPT phrase, and Izquierdo-Condoy et al. 2026 document stock LLM phrases appearing in papers that passed peer review.

    Supported1 sources · 1 papersTrace →
  9. 09

    Both Binz et al. 2025 and Walters & Wilder 2023 caution that verification costs may erode LLM efficiency gains: the trust routinely placed in ordinary software is inappropriate for generative AI, and time saved in text generation may be offset by the time required to verify the output.

    Supported2 sources · 2 papersTrace →
  10. 10

    Arzilli et al. 2025 warn that increased reliance on AI-driven authoring tools could exacerbate an infodemic of unreliable or misleading information, a risk amplified by the prevailing publish-or-perish culture.

    Supported2 sources · 1 papersTrace →
  11. 11

    Both Binz et al. 2025 and Walters & Wilder 2023 caution that verification costs may erode LLM efficiency gains: the trust routinely placed in ordinary software is inappropriate for generative AI, and time saved in text generation may be offset by the time required to verify the output.

    Not supported2 sources · 2 papersTrace →
  12. 12

    Independent detection methods converge in showing that LLM-generated text has rapidly entered the biomedical literature: Kobak et al. 2025's excess-vocabulary analysis of over 15 million PubMed abstracts and Fortenbach et al. 2026's detector-based screening of ophthalmology journals both document steep increases after ChatGPT's release.

    Supported2 sources · 2 papersTrace →
  13. 13

    Although AI-assisted writing is now common, disclosure of it is nearly nonexistent: He & Bu 2026 found almost no explicit disclosures in 75,000 recent papers, and Fortenbach et al. 2026 found none at all among ophthalmology publications flagged as likely AI-written.

    Supported2 sources · 2 papersTrace →
  14. 14

    Journal AI policies have so far been ineffective at changing author behavior: He & Bu 2026 show that despite 70% of journals adopting AI policies, mostly requiring disclosure, AI writing tool use grew just as much in journals with policies as in those without.

    Supported1 sources · 1 papersTrace →
  15. 15

    Walters & Wilder 2023 document that ChatGPT systematically fabricates bibliographic citations, with the systematic studies they review typically finding fabrication rates between 47% and 69%.

    Supported1 sources · 1 papersTrace →
  16. 16

    Bietti & Bangerter 2026 and Messeri & Crockett 2024 articulate a shared epistemic worry: because scientific writing is itself a form of thinking, outsourcing it to LLMs risks a mode of science that produces more text while understanding less.

    Supported2 sources · 2 papersTrace →
  17. 17

    Because individual cases of AI use in papers have proven impossible to adjudicate, researchers have turned to corpus-level statistical fingerprints: She 2026 frames this necessity, and Kobak et al. 2025's method implements it by detecting emerging LLM fingerprints directly from published abstracts rather than from ground-truth datasets.

    Supported2 sources · 2 papersTrace →
  18. 18

    LLM use has also entered peer review, not just manuscript writing: Lee et al. 2025 report corpus-level estimates of LLM-modified AI conference reviews, and She 2026 notes that such usage spikes as deadlines approach, revealing behavioral pressures behind adoption.

    Supported2 sources · 2 papersTrace →
  19. 19

    The democratizing promise of LLMs for non-native English writers, acknowledged even by critics such as Bietti & Bangerter 2026, is consistent with He & Bu 2026's large-scale finding that non-English-speaking countries exhibit the highest growth rates in AI-assisted writing.

    Supported2 sources · 2 papersTrace →
  20. 20

    Arzilli et al. 2025 and Bechky & Davis 2024 both warn that generative AI will inflate publication volume and further strain an already overcapacity publishing system, compounding publish-or-perish pressures and the risk of an unreliable-information infodemic.

    Partly supported1 sources · 1 papersTrace →
  21. 21

    Binz et al. 2025 and Walters & Wilder 2023 converge on the point that generative AI cannot be trusted the way conventional research software is: the verification burden falls on authors and may offset whatever time text generation saves.

    Supported2 sources · 2 papersTrace →
  22. 22

    Kobak et al. 2025's excess-vocabulary analysis of more than 15 million PubMed-indexed biomedical abstracts (2010–2024) suggests that at least 13.5% of 2024 abstracts were processed with LLMs, a lower bound reaching 40% for some subcorpora; the authors judge the resulting stylistic shift unprecedented, surpassing even the effect of the COVID pandemic.

    Supported3 sources · 1 papersTrace →
  23. 23

    Walters & Wilder 2023 show that fabricated references are a systematic failure mode of ChatGPT-assisted writing: across studies the share of fabricated citations typically falls in the 47–69% range, and one radiology evaluation found 64% of 343 citations could not be found in PubMed or on the open web.

    Supported2 sources · 1 papersTrace →
  24. 24

    He & Bu 2026 conclude that journal AI policies have largely failed: analyzing 5,114 journals and over 5.2 million papers, they find 70% of journals adopted policies (primarily requiring disclosure), yet AI writing tool use rose with no significant difference between journals with and without policies, and only about 0.1% of papers published since 2023 disclosed AI use.

    Supported3 sources · 1 papersTrace →
  25. 25

    Fortenbach et al. 2026 find that LLM-generated text is now common in ophthalmology: by 2025, 25.7% of sampled research articles and 21.6% of commentary articles carried AI-likelihood scores more than two standard deviations above baseline, and none of these outlier publications disclosed AI use.

    Supported2 sources · 1 papersTrace →
  26. 26

    LLM penetration varies sharply across fields at every stage of research: Lee et al. 2025 report corpus-level detection of 6.5–16.9% LLM-modified peer reviews at AI conferences but no significant evidence of LLM-based modifications in Nature journals, mirroring Kobak et al. 2025's observation that LLM-writing prevalence differs across disciplines, countries, and journals.

    Partly supported2 sources · 2 papersTrace →
  27. 27

    Fortenbach et al. 2026 transport Kobak et al. 2025's excess-vocabulary method into a single clinical specialty, screening ophthalmology abstracts for changes in word-frequency usage focused on stylistic words previously associated with LLM-generated text — the same abrupt style-word increases Kobak et al. used to date LLM adoption across 15 million biomedical abstracts.

    Partly supported2 sources · 2 papersTrace →
  28. 28

    Independent corpora converge on an AI transparency gap: He & Bu 2026 find that among 75,000 papers published since 2023 only 76 (~0.1%) disclosed AI use, while Fortenbach et al. 2026 report that none of the ophthalmology publications with outlier AI-likelihood scores disclosed their AI use.

    Supported2 sources · 2 papersTrace →
  29. 29

    Bietti & Bangerter 2026 anchor their warning that LLMs risk decoupling writing from thinking in prevalence estimates — more than one in ten biomedical papers published in 2024 and over 20% in some fields — consistent with the lower bounds Kobak et al. 2025 derived from excess vocabulary, showing how measurement studies now feed the normative debate.

    Supported2 sources · 2 papersTrace →
  30. 30

    Two perspective pieces converge on the epistemic stakes of AI-assisted science: Messeri & Crockett 2024 warn that proliferating AI tools risk a phase of enquiry in which we produce more but understand less, while Bietti & Bangerter 2026 argue that outsourcing writing to LLMs risks decoupling writing from thinking.

    Supported2 sources · 2 papersTrace →
  31. 31

    Bechky & Davis 2024, writing from a sociology-of-science standpoint, illustrate undisclosed LLM writing with an article that opened "Certainly, here is a possible introduction to your topic" and point readers to Kobak and colleagues' excess-vocabulary research for a systematic assessment in the medical literature, showing the measurement approach had already crossed disciplinary boundaries.

    Partly supported1 sources · 1 papersTrace →
  32. 32

    Fortenbach et al. 2026 find that LLM-generated text is now common in ophthalmology: by 2025, 25.7% of sampled research articles and 21.6% of commentary articles carried AI-likelihood scores more than two standard deviations above baseline, and none of these outlier publications disclosed AI use.

    Not supported2 sources · 2 papersTrace →
  33. 33

    Kobak et al. 2025 estimate that at least 13.5% of 2024 biomedical abstracts were processed with LLMs, with the lower bound reaching 40% for some subcorpora; they describe the resulting impact on scientific writing as unprecedented, surpassing the effect of major world events such as the COVID pandemic.

    Supported2 sources · 1 papersTrace →
  34. 34

    Walters & Wilder 2023 report that the proportion of fabricated citations in ChatGPT-generated content typically falls in the 47–69% range, citing a radiology study in which 64% of 343 citations were fabricated, i.e., could not be found in PubMed or on the open web.

    Supported1 sources · 1 papersTrace →
  35. 35

    He & Bu 2026 find that despite 70% of journals adopting AI policies, researchers' use of AI writing tools has increased dramatically across disciplines, with no significant difference between journals with or without policies.

    Supported1 sources · 1 papersTrace →
  36. 36

    Fortenbach et al. 2026 found that by 2025, 25.7% of sampled ophthalmology research articles and 21.6% of commentary articles contained AI-likelihood scores of more than 2 standard deviations above the baseline. Among publications with outlier scores, 22.3% of sentences in research articles and 90% of sentences in commentary articles were likely written by AI.

    Supported2 sources · 1 papersTrace →
  37. 37

    Kobak et al. 2025 and Fortenbach et al. 2026 both detect LLM-assisted writing through style-word frequencies: Kobak et al. studied vocabulary changes in more than 15 million PubMed abstracts, while Fortenbach et al. evaluated abstract text from 27,142 ophthalmology research articles for changes in word-frequency usage with a focus on stylistic words.

    Supported2 sources · 2 papersTrace →
  38. 38

    Non-disclosure emerges across studies using different methods: He & Bu 2026 found that of 75,000 papers published since 2023, only 76 (~0.1%) explicitly disclosed AI use, while Fortenbach et al. 2026 found that AI use was not disclosed among any ophthalmology publications with outlier AI-likelihood scores.

    Supported2 sources · 2 papersTrace →
  39. 39

    Bietti & Bangerter 2026 report that a recent study estimated more than one in 10 biomedical papers published in 2024 contained AI-generated text, often without disclosure, consistent with Kobak et al. 2025's lower bound of at least 13.5% of 2024 abstracts. They argue that using such tools can risk decoupling writing from thinking.

    Supported3 sources · 2 papersTrace →
  40. 40

    Lee et al. 2025 report LLM modification rates of 6.5% to 16.9% in AI conference peer reviews but no significant evidence of LLM-based modifications in Nature journals, a discipline-level variation that parallels Kobak et al. 2025's finding that LLM usage lower bounds differ across disciplines, countries, and journals.

    Supported2 sources · 2 papersTrace →
  41. 41

    Messeri & Crockett 2024 warn that the proliferation of AI tools in science risks a phase of scientific enquiry in which we produce more but understand less, anticipating Bietti & Bangerter 2026's contention that scientific writing is not merely communication but a form of thinking.

    Supported2 sources · 2 papersTrace →
  42. 42

    The language-equity rationale for LLM-assisted writing recurs across papers: Kobak et al. 2025 note that LLMs can help translate to English, Bietti & Bangerter 2026 note that LLMs can support non-native English speakers by reducing linguistic barriers, and He & Bu 2026 find that non-English-speaking countries exhibit the highest growth rates in AI writing tool use.

    Supported3 sources · 3 papersTrace →
  43. 43

    Walters & Wilder 2023 report that the proportion of fabricated citations in ChatGPT-generated content typically falls in the 47–69% range, citing a radiology study in which 64% of 343 citations were fabricated, i.e., could not be found in PubMed or on the open web.

    Not supported2 sources · 1 papersTrace →
  44. 44

    Kobak et al. 2025 measured LLM use in biomedical writing at an unprecedented scale, analyzing over 15 million PubMed abstracts from 2010–2024 and estimating that at least 13.5% of 2024 abstracts were processed with LLMs, with the lower bound reaching 40% in some subcorpora.

    Supported1 sources · 1 papersTrace →
  45. 45

    Kobak et al. 2025 argue their method is conceptually stronger than prior detection work because it detects LLM fingerprints directly from published abstracts rather than relying on potentially biased ground-truth datasets, allowing them to conclude that LLMs' impact on scientific writing surpasses even that of the COVID pandemic.

    Supported2 sources · 1 papersTrace →
  46. 46

    Independent detection approaches in different medical fields converge on the same pattern: the corpus-wide excess-vocabulary method of Kobak et al. 2025 and the journal-level AI-detection screening of Fortenbach et al. 2026 both document a sharp post-ChatGPT rise in LLM-generated text, with Fortenbach et al. finding over a quarter of sampled ophthalmology research articles showing outlier AI-likelihood scores by 2025.

    Supported2 sources · 2 papersTrace →
  47. 47

    Two large studies independently document a near-total transparency gap: He & Bu 2026 found only about 0.1% of 75,000 post-2023 papers explicitly disclosed AI use, and Fortenbach et al. 2026 found no disclosure in any ophthalmology publication flagged with outlier AI scores.

    Supported2 sources · 2 papersTrace →
  48. 48

    He & Bu 2026 show that journal AI policies have so far been ineffective: although 70% of 5,114 analyzed journals adopted AI policies, mostly requiring disclosure, AI writing-tool use grew dramatically with no significant difference between journals with and without such policies.

    Supported1 sources · 1 papersTrace →
  49. 49

    Walters & Wilder 2023 document that fabricated citations are a systematic failure mode of ChatGPT, with fabrication rates typically between 47% and 69% across studies, and argue that the trust researchers reasonably place in statistical software is not warranted for generative AI because its tasks are fundamentally different.

    Supported2 sources · 1 papersTrace →
  50. 50

    Messeri & Crockett 2024 and Bietti & Bangerter 2026 raise complementary epistemic alarms about AI in science: the former warn of a phase of enquiry in which we produce more but understand less, while the latter argue that outsourcing writing to LLMs risks decoupling writing from thinking, since scientific writing is itself a form of thinking.

    Supported3 sources · 2 papersTrace →

Anchor papers

B
PaperYearLicense
Quantifying large language model usage in scientific papers
Nature Human Behaviour
A measurable rise in LLM use across fields.
2025Red
Delving into LLM-assisted writing in biomedical publications through excess vocabulary
Science Advances
At least 13.5% of 2024 PubMed abstracts were processed with LLMs.
2025Green
A Critical Analysis of Generative AI: Challenges, Opportunities, and Future Research Directions
Archives of Computational Methods in Engineering
Added by Fredrik in Zotero, 2026-09-30: critical analysis of generative AI.
2025Green
RETRACTED: The three-dimensional porous mesh structure of Cu-based metal-organic-framework - Aramid cellulose separator enhances the electrochemical performance of lithium metal anode batteries
Surfaces and Interfaces Retracted
Retracted 2024 (Retraction Watch 53671): the introduction opened with "Certainly, here is". Publisher terms only: red and retracted.
2024Red
Artificial intelligence and illusions of understanding in scientific research
Nature
Illusions of understanding. The load-bearing text for act 2.
2024Red
Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
arXiv (Cornell University)· preprint
LLM-generated text in peer review.
2024Orange
RETRACTED: Cellular functions of spermatogonial stem cells in relation to JAK/STAT signaling pathway
Frontiers in Cell and Developmental Biology Retracted
Retracted 2024 (Retraction Watch 51983): AI-generated figures. CC-BY yet retracted: the license says it may be stored, the retraction keeps it out of synthesis.
2024Green
Fabrication and errors in the bibliographic citations generated by ChatGPT
Scientific Reports
55% fabricated references with GPT-3.5, 18% with GPT-4.
2023Green
AI and science: what 1,600 researchers think
Nature
Researchers' own concerns.
2023Red

Found through citations

C
PaperYearLicense
Will the widespread use of large language models in scientific writing undermine scientists’ critical thinking?
PLoS Biology
discovered: cites 2 anchor(s) (Delving into LLM-assisted writing in biomedical pu; Quantifying large language model usage in scientif); similarity 0.75
2026Green
Artificial Intelligence for Academic Text Generation in Analytical Chemistry: Current Risks, Indicators, and Perspectives toward Greener and More Sustainable Approaches
Analytical Chemistry
discovered: cites 2 anchor(s) (Delving into LLM-assisted writing in biomedical pu; Monitoring AI-Modified Content at Scale: A Case St); similarity 0.79
2026Red
Academic journals’ AI policies fail to curb the surge in AI-assisted academic writing
Proceedings of the National Academy of Sciences
discovered: cites 3 anchor(s) (Delving into LLM-assisted writing in biomedical pu; Quantifying large language model usage in scientif; AI and science: what 1,600 researchers think); similarity 0.73
2026Yellow
Large Language Model Authorship in Ophthalmic Publications
Ophthalmology
discovered: cites 2 anchor(s) (Delving into LLM-assisted writing in biomedical pu; Quantifying large language model usage in scientif); similarity 0.75
2026Yellow
Fine-Grained Detection of AI-Generated Writing in the Biomedical Literature
bioRxiv (Cold Spring Harbor Laboratory)· preprint
discovered: cites 2 anchor(s) (Monitoring AI-Modified Content at Scale: A Case St; Delving into LLM-assisted writing in biomedical pu); similarity 0.73
2026Green
Artificial Intelligence in Medical Education: Transformative Potential, Current Applications, and Future Implications
JMIR Medical Education
discovered: cites 1 anchor(s) (Fabrication and errors in the bibliographic citati); similarity 0.77
2026Green
The Epistemic Downside of Using LLM-Based Generative AI in Academic Writing
Publications
discovered: cites 2 anchor(s) (Delving into LLM-assisted writing in biomedical pu; Quantifying large language model usage in scientif); similarity 0.75
2025Green
A Primer for Evaluating Large Language Models in Social-Science Research
Advances in Methods and Practices in Psychological Science
discovered: cites 2 anchor(s) (Artificial intelligence and illusions of understan; Delving into LLM-assisted writing in biomedical pu); similarity 0.75
2025Yellow
The role of large language models in the peer-review process: opportunities and challenges for medical journal reviewers and editors
Journal of Educational Evaluation for Health Professions
discovered: cites 2 anchor(s) (Monitoring AI-Modified Content at Scale: A Case St; AI and science: what 1,600 researchers think); similarity 0.77
2025Green
Generative Artificial Intelligence (GenAI) in the research process – A survey of researchers’ practices and perceptions
Technology in Society
discovered: cites 2 anchor(s) (Artificial intelligence and illusions of understan; AI and science: what 1,600 researchers think); similarity 0.75
2025Green
How should the advancement of large language models affect the practice of science?
Proceedings of the National Academy of Sciences
discovered: cites 2 anchor(s) (Fabrication and errors in the bibliographic citati; Artificial intelligence and illusions of understan); similarity 0.73
2025Green
Generative Artificial Intelligence: Implications for Biomedical and Health Professions Education
Annual Review of Biomedical Data Science
discovered: cites 1 anchor(s) (Fabrication and errors in the bibliographic citati); similarity 0.77
2025Green
A surge of AI-driven publications: the impact on health professionals and potential mitigating solutions
Frontiers in Public Health
discovered: cites 1 anchor(s) (RETRACTED: Cellular functions of spermatogonial st); similarity 0.76
2025Green
The role of generative AI in academic and scientific authorship: an autopoietic perspective
AI & Society
discovered: cites 2 anchor(s) (Monitoring AI-Modified Content at Scale: A Case St; RETRACTED: Cellular functions of spermatogonial st); similarity 0.77
2025Green
The Use of Large Language Models and Their Association With Enhanced Impact in Biomedical Research and Beyond
MedComm – Future Medicine
discovered: cites 1 anchor(s) (Monitoring AI-Modified Content at Scale: A Case St); similarity 0.76
2025Green
“AI et al.” The perils of overreliance on Artificial Intelligence by authors in scientific research
Clinical eHealth
discovered: cites 2 anchor(s) (RETRACTED: Cellular functions of spermatogonial st; RETRACTED: The three-dimensional porous mesh struc); similarity 0.77
2024Yellow
ChatGPT in Teaching and Learning: A Systematic Review
Education Sciences
discovered: cites 1 anchor(s) (Fabrication and errors in the bibliographic citati); similarity 0.76
2024Green
Evaluating Literature Reviews Conducted by Humans Versus ChatGPT: Comparative Study
JMIR AI
discovered: cites 1 anchor(s) (AI and science: what 1,600 researchers think); similarity 0.77
2024Green
The transformative impact of large language models on medical writing and publishing: current applications, challenges and future directions
Korean Journal of Physiology and Pharmacology
discovered: cites 1 anchor(s) (Monitoring AI-Modified Content at Scale: A Case St); similarity 0.77
2024Yellow
The ethics of using artificial intelligence in scientific research: new guidance needed for a new tool
AI and Ethics
discovered: cites 2 anchor(s) (Artificial intelligence and illusions of understan; Fabrication and errors in the bibliographic citati); similarity 0.75
2024Green
ChatGPT and generative AI are revolutionizing the scientific community: A Janus‐faced conundrum
iMeta
discovered: cites 1 anchor(s) (AI and science: what 1,600 researchers think); similarity 0.77
2024Green
Resisting the Algorithmic Management of Science: Craft and Community After Generative AI
Administrative Science Quarterly
discovered: cites 3 anchor(s) (Artificial intelligence and illusions of understan; Delving into LLM-assisted writing in biomedical pu; Monitoring AI-Modified Content at Scale: A Case St); similarity 0.69
2024Green
conflict
Artificial Intelligence: A Challenge to Scientific Communication
Klinische Monatsblätter für Augenheilkunde
discovered: cites 1 anchor(s) (RETRACTED: Cellular functions of spermatogonial st); similarity 0.79
2024Red
Obvious artificial intelligence ‐generated anomalies in published journal articles: A call for enhanced editorial diligence
Learned Publishing
discovered: cites 2 anchor(s) (RETRACTED: Cellular functions of spermatogonial st; RETRACTED: The three-dimensional porous mesh struc); similarity 0.74
2024Green
Implementation and Evaluation of a ChatGPT-Assisted Special Topics Writing Assignment in Biochemistry
Journal of Chemical Education
discovered: cites 2 anchor(s) (Fabrication and errors in the bibliographic citati; AI and science: what 1,600 researchers think); similarity 0.73
2024Red