Because individual cases of AI use in papers have proven impossible to adjudicate, researchers have turned to corpus-level statistical fingerprints: She 2026 frames this necessity, and Kobak et al. 2025's method implements it by detecting emerging LLM fingerprints directly from published abstracts rather than from ground-truth datasets.
She states that individual cases of AI usage have proven impossible to adjudicate, motivating coarse-grained corpus-level metrics, and Kobak et al. describe detecting emerging LLM fingerprints directly from published abstracts rather than ground-truth datasets.
Written by Kimi K3 via Ollama Cloud · checked by GLM-5.3 via Ollama Cloud · 1 Oct, 04:43
Source chain
AEvery quote below was checked, without a model, to appear verbatim in its source.
- 01
“Since individual cases of AI usage have thus far proven impossible to adjudicate, existing studies have focused on coarse grained metrics.”
Full-text passage · p. 2 · evidence E22
of these systems, one might reasonably wonder to what extent AI-generated writing has already made its way into the scientific literature. Since individual cases of AI usage have thus far proven impossible to adjudicate, existing studies have focused on coarse grained metrics. One study that analyzed over 15 million abstracts from the biomedical literature documented the emergence of a bona fide LLM lexicon, with specific model-favored words rising at rates that cannot be explained by natural linguistic drift2. Other case reports point to the use of generative AI during peer review3, suggesting that academics find these tools both useful and embarrassing, deploying them only behind the shield of anonymity for work that never becomes part of the permanent record. A broader analysis of conference peer-review data reveals that the estimat ed fraction of LLM-generated reviews spikes as deadlines approach4, underscoring the messy practical and behavioral variables that determine where and when generative AI is used. However, these studies leave several critical questions unanswered: 1) who exactly is using generative AI?
Passage read from www.biorxiv.org, which may be a preprint rather than the published version.
- 02
“our analysis avoids this limitation by detecting emerging LLM fingerprints directly from published abstracts.”
Full-text passage · no page number · evidence E6
Passage read from www.ebi.ac.uk, which may be a preprint rather than the published version.