How should the advancement of large language models affect the practice of science?
Marcel Binz, Stephan Alaniz, Adina L. Roskies, Balázs Aczél, Carl T. Bergstrom, Colin Allen, Daniel J. Schad, Dirk U. Wulff and 10 more
Why it has this license class
AChecked 1 Oct 2026. Open license (CC-BY, CC-BY-SA, CC0, public domain): full text indexed and used in synthesis.
| Source | License | Open-access status | Read as |
|---|---|---|---|
| openalex | cc-by | hybrid | Green |
| crossref | https://creativecommons.org/licenses/by/4.0/ | — | Green |
| unpaywall | cc-by | hybrid | Green |
| europepmc | cc by | — | Green |
Abstract
BLarge language models (LLMs) are being increasingly incorporated into scientific workflows. However, we have yet to fully grasp the implications of this integration. How should the advancement of large language models affect the practice of science? For this opinion piece, we have invited four diverse groups of scientists to reflect on this query, sharing their perspectives and engaging in debate. Schulz et al. make the argument that working with LLMs is not fundamentally different from working with human collaborators, while Bender et al. argue that LLMs are often misused and overhyped, and that their limitations warrant a focus on more specialized, easily interpretable tools. Marelli et al. emphasize the importance of transparent attribution and responsible use of LLMs. Finally, Botvinick and Gershman advocate that humans should retain responsibility for determining the scientific roadmap. To facilitate the discussion, the four perspectives are complemented with a response from each group. By putting these different perspectives in conversation, we aim to bring attention to important considerations within the academic community regarding the adoption of LLMs and their impact on both current and future scientific practices.
Claims built on this paper
D- Binz et al. 2025 and Walters & Wilder 2023 converge on the point that generative AI cannot be trusted the way conventional research software is: the verification burden falls on authors and may offset whatever time text generation saves. Trace →
- Both Binz et al. 2025 and Walters & Wilder 2023 caution that verification costs may erode LLM efficiency gains: the trust routinely placed in ordinary software is inappropriate for generative AI, and time saved in text generation may be offset by the time required to verify the output. Trace →