Large Language Model Authorship in Ophthalmic Publications
Christopher R. Fortenbach, Yue S. Wu, Parth M. Mungra, Russell Neil Van Gelder
Why it has this license class
AChecked 30 Sept 2026. Non-commercial or no-derivatives license: full text kept internally; only metadata and abstract are indexed and used.
| Source | License | Open-access status | Read as |
|---|---|---|---|
| openalex | cc-by-nc-nd | hybrid | Yellow |
| crossref | http://creativecommons.org/licenses/by-nc-nd/4.0/ | — | Yellow |
| unpaywall | cc-by-nc-nd | hybrid | Yellow |
Abstract
BPURPOSE: To assess for the likely presence of artificial intelligence (AI)-generated text in the published ophthalmology literature. METHODS: Abstract text from 27 142 research articles published in 22 journals between May 2020 and May 2025 was evaluated for changes in word-frequency usage with a focus on stylistic words previously found to be associated with large language model (LLM)-generated text. Four commercial AI-detection services (ZeroGPT, Writer.com, Winston AI, GPTZero) were first validated against control articles with GPTZero showing the best performance, which was then used to detect the presence of AI-generated text in 50 full articles from each journal. For the large-scale screening, research articles and commentary publications (e.g., editorials) were scored at the section and sentence level and compared in the pre- versus post-ChatGPT publication time periods. RESULTS: Since the release of ChatGPT in 2022, a marked increase in previously rarely used stylistic words was observed with at least a 2-fold usage increase observed in 20% of ophthalmology abstracts. With full article text evaluation, GPTZero scores increased after the release of ChatGPT across all research article sections (e.g., abstract, introduction) and commentary articles. By 2025, 25.7% of sampled research articles and 21.6% of commentary articles contained AI-likelihood scores of more than 2 standard deviations above the baseline. Sentence-level analysis showed that among those publications containing outlier scores, 22.3% of sentences in research articles and 90% of sentences in commentary articles were likely written by AI. Use of AI was not disclosed among any of the publications with outlier scores. CONCLUSIONS: Artificial intelligence brings significant promise in its ability to facilitate both scientific and medical advances. As these tools become more powerful, disclosure regarding the manner of their use becomes increasingly important. We show that LLM-generated text is increasingly present in the ophthalmic literature and is rarely disclosed. Without disclosure requirements and editorial oversight, there is a significant risk that undisclosed LLM usage will continue to increase and may jeopardize authorship integrity and long-term reliability of published findings. FINANCIAL DISCLOSURE(S): Proprietary or commercial disclosure may be found after the references.
Claims built on this paper
D- Fortenbach et al. (2025/2026) find the same kind of stylistic-word surge in ophthalmology that Kobak et al. (2025) found in biomedicine, and additionally use AI detectors to report that 25.7% of sampled research articles had outlier AI-likelihood scores by 2025. Trace →
- Disclosure of LLM use is rare: Fortenbach et al. (2026) found no disclosure among ophthalmic publications with outlier scores, and He & Bu (2026) found only 76 of 75 thousand post-2023 papers explicitly disclosed AI use. Trace →
- Walters & Wilder (2023) argue that trust in conventional software does not carry over to generative AI, which matters given the evidence of undisclosed LLM use in published writing (Fortenbach et al. 2026). Trace →
- Two corpus-level studies measure LLM use through surges in stylistic words. Kobak et al. 2025 estimate that at least 13.5% of 2024 PubMed abstracts were LLM-processed, and Fortenbach et al. 2026 find at least a 2-fold usage increase in 20% of ophthalmology abstracts after ChatGPT's release. Trace →
- Across independent datasets, AI-assisted writing is almost never disclosed. He & Bu 2026 find that only 76 of 75k post-2023 papers (~0.1%) disclosed AI use. Fortenbach et al. 2026 find no disclosure among any ophthalmology publications with outlier AI-likelihood scores. Trace →
- Walters & Wilder 2023 report that the fabricated share of ChatGPT-generated citations typically falls in the 47–69% range. They argue that generative AI output does not merit the unchecked trust routinely given to other research software. Trace →
- Independent detection approaches in different medical fields converge on the same pattern: the corpus-wide excess-vocabulary method of Kobak et al. 2025 and the journal-level AI-detection screening of Fortenbach et al. 2026 both document a sharp post-ChatGPT rise in LLM-generated text, with Fortenbach et al. finding over a quarter of sampled ophthalmology research articles showing outlier AI-likelihood scores by 2025. Trace →
- Two large studies independently document a near-total transparency gap: He & Bu 2026 found only about 0.1% of 75,000 post-2023 papers explicitly disclosed AI use, and Fortenbach et al. 2026 found no disclosure in any ophthalmology publication flagged with outlier AI scores. Trace →
- Kobak et al. 2025 measured LLM use in biomedical writing at an unprecedented scale, analyzing over 15 million PubMed abstracts from 2010–2024 and estimating that at least 13.5% of 2024 abstracts were processed with LLMs, with the lower bound reaching 40% in some subcorpora. Trace →
- Fortenbach et al. 2026 found that by 2025, 25.7% of sampled ophthalmology research articles and 21.6% of commentary articles contained AI-likelihood scores of more than 2 standard deviations above the baseline. Among publications with outlier scores, 22.3% of sentences in research articles and 90% of sentences in commentary articles were likely written by AI. Trace →
- Kobak et al. 2025 and Fortenbach et al. 2026 both detect LLM-assisted writing through style-word frequencies: Kobak et al. studied vocabulary changes in more than 15 million PubMed abstracts, while Fortenbach et al. evaluated abstract text from 27,142 ophthalmology research articles for changes in word-frequency usage with a focus on stylistic words. Trace →
- Non-disclosure emerges across studies using different methods: He & Bu 2026 found that of 75,000 papers published since 2023, only 76 (~0.1%) explicitly disclosed AI use, while Fortenbach et al. 2026 found that AI use was not disclosed among any ophthalmology publications with outlier AI-likelihood scores. Trace →
- Walters & Wilder 2023 report that the proportion of fabricated citations in ChatGPT-generated content typically falls in the 47–69% range, citing a radiology study in which 64% of 343 citations were fabricated, i.e., could not be found in PubMed or on the open web. Trace →
- Fortenbach et al. 2026 find that LLM-generated text is now common in ophthalmology: by 2025, 25.7% of sampled research articles and 21.6% of commentary articles carried AI-likelihood scores more than two standard deviations above baseline, and none of these outlier publications disclosed AI use. Trace →
- Fortenbach et al. 2026 transport Kobak et al. 2025's excess-vocabulary method into a single clinical specialty, screening ophthalmology abstracts for changes in word-frequency usage focused on stylistic words previously associated with LLM-generated text — the same abrupt style-word increases Kobak et al. used to date LLM adoption across 15 million biomedical abstracts. Trace →
- Independent corpora converge on an AI transparency gap: He & Bu 2026 find that among 75,000 papers published since 2023 only 76 (~0.1%) disclosed AI use, while Fortenbach et al. 2026 report that none of the ophthalmology publications with outlier AI-likelihood scores disclosed their AI use. Trace →
- Independent detection methods converge in showing that LLM-generated text has rapidly entered the biomedical literature: Kobak et al. 2025's excess-vocabulary analysis of over 15 million PubMed abstracts and Fortenbach et al. 2026's detector-based screening of ophthalmology journals both document steep increases after ChatGPT's release. Trace →
- Although AI-assisted writing is now common, disclosure of it is nearly nonexistent: He & Bu 2026 found almost no explicit disclosures in 75,000 recent papers, and Fortenbach et al. 2026 found none at all among ophthalmology publications flagged as likely AI-written. Trace →
- Field-specific screening corroborates the biomedical-wide estimate: Fortenbach et al. 2026 found that by 2025 about a quarter of sampled ophthalmology research articles carried AI-likelihood scores more than two standard deviations above baseline, consistent with the lower bound reported by Kobak et al. 2025. Trace →
- Disclosure is almost entirely absent even where LLM use is detected: He & Bu 2026 found only about 0.1% of post-2023 papers disclosed AI use despite most journals having policies, and Fortenbach et al. 2026 found no disclosure among any publications with outlier scores, indicating current policies fail to promote transparency. Trace →