Claim 26 · research-integrity · finding

Walters & Wilder 2023 document that fabricated citations are a systematic failure mode of ChatGPT, with fabrication rates typically between 47% and 69% across studies, and argue that the trust researchers reasonably place in statistical software is not warranted for generative AI because its tasks are fundamentally different.

Supported

E11 states fabricated citations are systematically investigated with rates typically in the 47–69% range, and E12 argues the justified trust placed in statistical software is not appropriate for generative AI because its tasks are fundamentally different.

Written by Kimi K3 via Ollama Cloud · 30 Sept, 22:48

Source chain

A

Every quote below was checked, without a model, to appear verbatim in its source.

  1. 01

    “the proportion of fabricated citations is typically in the 47–69% range, with a higher rate in geography than in medicine”

    Full-text passage · no page number · evidence E11

    Passage read from www.ebi.ac.uk, which may be a preprint rather than the published version.

  2. 02

    “That same level of trust is not appropriate with generative AI tools, however, since the tasks performed by AI are fundamentally different”

    Full-text passage · no page number · evidence E12

    Passage read from www.ebi.ac.uk, which may be a preprint rather than the published version.