AI capability is uneven across scientific task types: Skarlinski et al. 2024 show PaperQA2 matches or exceeds experts on literature research tasks, yet Alampara et al. 2025 find vision language models fail at spatial reasoning and cross-modal synthesis, and Minasny et al. 2026 report LLMs answer only up to 65% of advanced soil science exam questions correctly.
E12 shows PaperQA2 matches or exceeds subject-matter experts on literature research tasks, E42 shows vision language models exhibit fundamental limitations in spatial reasoning and cross-modal information synthesis, and E23 reports LLMs correctly answer only up to 65% of advanced soil science exam questions.
Written by Kimi K3 via Ollama Cloud · checked by GLM-5.3 via Ollama Cloud · 1 Oct, 04:32
Source chain
AEvery quote below was checked, without a model, to appear verbatim in its source.
- 01
“matches or exceeds subject matter expert performance on three realistic literature research tasks without any restrictions on humans”
Full-text passage · p. 1 · evidence E12
LANGUAGE AGENTS ACHIEVE SUPERHUMAN SYNTHESIS OF SCIENTIFIC KNOWLEDGE Michael D. Skarlinski1 Sam Cox1,2 Jon M. Laurent1 James D. Braza1 Michaela Hinks1 Michael J. Hammerling1 Manvitha Ponnapati1 Samuel G. Rodriques1,3∗ Andrew D. White1,2∗ 1FutureHouse Inc., San Francisco, CA 2University of Rochester, Rochester, NY 3 Francis Crick Institute, London, UK ∗These authors jointly supervise technical work at FutureHouse. Correspondence to: {sam,andrew}@futurehouse.org ABSTRACT Language models are known to “hallucinate” incorrect information, and it is unclear if they are sufficiently accurate and reliable for use in scientific research. We developed a rigorous human-AI comparison methodology to evaluate language model agents on real-world literature search tasks covering information retrieval, summarization, and contradiction detection tasks. We show that PaperQA2, a frontier language model agent optimized for improved factuality, matches or exceeds subject matter expert performance on three realistic literature research tasks without any restrictions on humans (i.e., full access to internet, search tools, and time).
Passage read from arxiv.org, which may be a preprint rather than the published version.
- 02
“they exhibit fundamental limitations in spatial reasoning, cross-modal information synthesis and multi-step logical inference”
Full-text passage · no page number · evidence E42
Passage read from www.ebi.ac.uk, which may be a preprint rather than the published version.
- 03
“on average, they could correctly answer only up to 65% of questions from advanced soil science examinations”
Full-text passage · p. 3 · evidence E23
While they demonstrate promising capabilities for conversational interaction, ensuring their reliability and depth of understanding in specialized soil science domains still requires signifi cant human validation and expertise. For example, Khanifar ( 34) assessed the performance of LLMs in answering soil science-related questions and found that, on average, they could correctly answer only up to 65% of questions from advanced soil science examinations. This existing foundation, with its strengths and weaknesses, sets the stage for the development of next-generation AI agents capable of more integrated, autonomou s, and collaborative scienti fic exploration and management in soil science. The vision of AI agents transforming scienti fic discovery is gaining momentum in fields such as chemistry, materials science, and biomedicine, where complex systems, vast datasets, and multidisciplinary knowledge must be integrated ( 5, 35, 36). The concept of an “AI scientist” or “co-scientist”— an AI system capable of skeptical learning, reasoning, and collaboration— has been proposed as a framework for generating novel scienti fich y p o t h e s e s aligned with researcher objectives ( 8, 37, 38).
Passage read from www.frontiersin.org, which may be a preprint rather than the published version.