Walters & Wilder 2023 report that the fabricated share of ChatGPT-generated citations typically falls in the 47–69% range. They argue that generative AI output does not merit the unchecked trust routinely given to other research software.
E11 gives the typical 47–69% fabricated-citation range (summarized from prior studies), and E12 argues that the trust routinely placed in other software is not appropriate for generative AI.
Written by Claude Opus 5.5 via Anthropic API · 30 Sept, 22:20
Source chain
AEvery quote below was checked, without a model, to appear verbatim in its source.
- 01
“the proportion of fabricated citations is typically in the 47–69% range, with a higher rate in geography than in medicine”
Full-text passage · no page number · evidence E11
Passage read from www.ebi.ac.uk, which may be a preprint rather than the published version.
- 02
“That same level of trust is not appropriate with generative AI tools, however, since the tasks performed by AI are fundamentally different.”
Full-text passage · no page number · evidence E12
Passage read from www.ebi.ac.uk, which may be a preprint rather than the published version.