Fabrication and errors in the bibliographic citations generated by ChatGPT
William H. Walters, Esther Isabelle Wilder
Why it has this license class
AChecked 30 Sept 2026. Open license (CC-BY, CC-BY-SA, CC0, public domain): full text indexed and used in synthesis.
| Source | License | Open-access status | Read as |
|---|---|---|---|
| openalex | cc-by | gold | Green |
| crossref | https://creativecommons.org/licenses/by/4.0 | — | Green |
| unpaywall | cc-by | gold | Green |
| europepmc | cc by | — | Green |
Abstract
BAlthough chatbots such as ChatGPT can facilitate cost-effective text generation and editing, factually incorrect responses (hallucinations) limit their utility. This study evaluates one particular type of hallucination: fabricated bibliographic citations that do not represent actual scholarly works. We used ChatGPT-3.5 and ChatGPT-4 to produce short literature reviews on 42 multidisciplinary topics, compiling data on the 636 bibliographic citations (references) found in the 84 papers. We then searched multiple databases and websites to determine the prevalence of fabricated citations, to identify errors in the citations to non-fabricated papers, and to evaluate adherence to APA citation format. Within this set of documents, 55% of the GPT-3.5 citations but just 18% of the GPT-4 citations are fabricated. Likewise, 43% of the real (non-fabricated) GPT-3.5 citations but just 24% of the real GPT-4 citations include substantive citation errors. Although GPT-4 is a major improvement over GPT-3.5, problems remain.
Claims built on this paper
D- Walters & Wilder (2023) summarize prior studies showing that ChatGPT often fabricates citations, with fabricated proportions typically 47–69%, including 64% of 343 citations in a radiology test. Trace →
- Walters & Wilder (2023) argue that trust in conventional software does not carry over to generative AI, which matters given the evidence of undisclosed LLM use in published writing (Fortenbach et al. 2026). Trace →
- Walters & Wilder 2023 report that the fabricated share of ChatGPT-generated citations typically falls in the 47–69% range. They argue that generative AI output does not merit the unchecked trust routinely given to other research software. Trace →
- Walters & Wilder 2023 document that fabricated citations are a systematic failure mode of ChatGPT, with fabrication rates typically between 47% and 69% across studies, and argue that the trust researchers reasonably place in statistical software is not warranted for generative AI because its tasks are fundamentally different. Trace →
- Walters & Wilder 2023 report that the proportion of fabricated citations in ChatGPT-generated content typically falls in the 47–69% range, citing a radiology study in which 64% of 343 citations were fabricated, i.e., could not be found in PubMed or on the open web. Trace →
- Walters & Wilder 2023 show that fabricated references are a systematic failure mode of ChatGPT-assisted writing: across studies the share of fabricated citations typically falls in the 47–69% range, and one radiology evaluation found 64% of 343 citations could not be found in PubMed or on the open web. Trace →
- Walters & Wilder 2023 document that ChatGPT systematically fabricates bibliographic citations, with the systematic studies they review typically finding fabrication rates between 47% and 69%. Trace →
- Binz et al. 2025 and Walters & Wilder 2023 converge on the point that generative AI cannot be trusted the way conventional research software is: the verification burden falls on authors and may offset whatever time text generation saves. Trace →
- Walters & Wilder 2023 document that ChatGPT fabricates a large share of the bibliographic citations it generates, with systematic studies typically finding fabrication rates in the 47–69% range. Trace →
- Both Binz et al. 2025 and Walters & Wilder 2023 caution that verification costs may erode LLM efficiency gains: the trust routinely placed in ordinary software is inappropriate for generative AI, and time saved in text generation may be offset by the time required to verify the output. Trace →