Language agents achieve superhuman synthesis of scientific knowledge
Michael Skarlinski, Sam Cox, Jon M. Laurent, James D. Braza, Michaela M. Hinks, Michael J. Hammerling, Manvitha Ponnapati, Samuel G. Rodriques and 1 more
Why it has this license class
AChecked 30 Sept 2026. Open license (CC-BY, CC-BY-SA, CC0, public domain): full text indexed and used in synthesis.
| Source | License | Open-access status | Read as |
|---|---|---|---|
| openalex | cc-by-sa | green | Green |
| arxiv | http://creativecommons.org/licenses/by-sa/4.0/ | — | Green |
Abstract
BLanguage models are known to hallucinate incorrect information, and it is unclear if they are sufficiently accurate and reliable for use in scientific research. We developed a rigorous human-AI comparison methodology to evaluate language model agents on real-world literature search tasks covering information retrieval, summarization, and contradiction detection tasks. We show that PaperQA2, a frontier language model agent optimized for improved factuality, matches or exceeds subject matter expert performance on three realistic literature research tasks without any restrictions on humans (i.e., full access to internet, search tools, and time). PaperQA2 writes cited, Wikipedia-style summaries of scientific topics that are significantly more accurate than existing, human-written Wikipedia articles. We also introduce a hard benchmark for scientific literature research called LitQA2 that guided design of PaperQA2, leading to it exceeding human performance. Finally, we apply PaperQA2 to identify contradictions within the scientific literature, an important scientific task that is challenging for humans. PaperQA2 identifies 2.34 +/- 1.99 contradictions per paper in a random subset of biology papers, of which 70% are validated by human experts. These results demonstrate that language model agents are now capable of exceeding domain experts across meaningful tasks on scientific literature.
Claims built on this paper
D- Evaluation infrastructure is being built alongside the agents themselves, with explicit human comparison as the yardstick: Skarlinski et al. 2024 report PaperQA2 matches or exceeds subject-matter experts on realistic literature tasks, and Laurent et al. 2024's LAB-Bench contributes over 2,400 biology research questions benchmarked against PhD-level scientists. Miller et al. 2025 identify the remaining gap—end-to-end biomedical ML workflows—and introduce BioML-bench to cover it. Trace →
- AI capability is uneven across scientific task types: Skarlinski et al. 2024 show PaperQA2 matches or exceeds experts on literature research tasks, yet Alampara et al. 2025 find vision language models fail at spatial reasoning and cross-modal synthesis, and Minasny et al. 2026 report LLMs answer only up to 65% of advanced soil science exam questions correctly. Trace →