Claim 59 · ai-for-science · connection

Evaluation infrastructure is being built alongside the agents themselves, with explicit human comparison as the yardstick: Skarlinski et al. 2024 report PaperQA2 matches or exceeds subject-matter experts on realistic literature tasks, and Laurent et al. 2024's LAB-Bench contributes over 2,400 biology research questions benchmarked against PhD-level scientists. Miller et al. 2025 identify the remaining gap—end-to-end biomedical ML workflows—and introduce BioML-bench to cover it.

Partly supported

The PaperQA2 and BioML-bench assertions are supported, but the cited LAB-Bench text says results were compared to 'human expert biology researchers' without specifying 'PhD-level scientists'.

Written by Kimi K3 via Ollama Cloud · checked by GLM-5.3 via Ollama Cloud · 30 Sept, 23:08

Source chain

A

Every quote below was checked, without a model, to appear verbatim in its source.

  1. 01

    “matches or exceeds subject matter expert performance on three realistic literature research tasks without any restrictions on humans”

    Full-text passage · p. 1 · evidence E15

    LANGUAGE AGENTS ACHIEVE SUPERHUMAN SYNTHESIS OF SCIENTIFIC KNOWLEDGE Michael D. Skarlinski1 Sam Cox1,2 Jon M. Laurent1 James D. Braza1 Michaela Hinks1 Michael J. Hammerling1 Manvitha Ponnapati1 Samuel G. Rodriques1,3∗ Andrew D. White1,2∗ 1FutureHouse Inc., San Francisco, CA 2University of Rochester, Rochester, NY 3 Francis Crick Institute, London, UK ∗These authors jointly supervise technical work at FutureHouse. Correspondence to: {sam,andrew}@futurehouse.org ABSTRACT Language models are known to “hallucinate” incorrect information, and it is unclear if they are sufficiently accurate and reliable for use in scientific research. We developed a rigorous human-AI comparison methodology to evaluate language model agents on real-world literature search tasks covering information retrieval, summarization, and contradiction detection tasks. We show that PaperQA2, a frontier language model agent optimized for improved factuality, matches or exceeds subject matter expert performance on three realistic literature research tasks without any restrictions on humans (i.e., full access to internet, search tools, and time).

    Passage read from arxiv.org, which may be a preprint rather than the published version.

  2. 02

    “over 2,400 multiple choice questions for evaluating AI systems on a range of practical biology research capabilities”

    Full-text passage · p. 1 · evidence E16

    As a step toward building such benchmarks, we introduce the Language Agent Biology Benchmark (LAB-Bench), a broad dataset of over 2,400 multiple choice questions for evaluating AI systems on a range of practical biology research capabilities, including recall and reasoning over literature, interpretation of figures, access and navigation of databases, and comprehension and manipulation of DNA and protein sequences. Importantly, in contrast to previous scientific benchmarks, we expect that an AI system that can achieve consistently high scores on the more difficult LAB-Bench tasks would serve as a useful assistant for researchers in areas such as literature search and molecular cloning. As an initial assessment of the emergent scientific task capabilities of frontier language models, we measure performance of several against our benchmark and report results compared to human expert biology researchers. We will continue to update and expand LAB-Bench over time, and expect it to serve as a useful tool in the development of automated research systems going forward.

    Passage read from arxiv.org, which may be a preprint rather than the published version.

  3. 03

    “the first benchmarking suite for evaluating AI agents on end-to-end biomedical ML tasks”

    Full-text passage · p. 1 · evidence E35

    BioML-bench: Evaluation of AI Agents for End-to-End Biomedical ML Henry E. Miller∗† Shift Bioscience Toronto, Canada Matthew Greenig∗ University of Cambridge ScienceMachine Cambridge, UK Benjamin Tenmann ScienceMachine London, UK Bo Wang University of Toronto Vector Institute Toronto, Canada Abstract Large language model (LLM) agents hold promise for accelerating biomedical research and development (R&D). Several biomedical agents have recently been proposed, but their evaluation has largely been restricted to question answering (e.g., LAB-Bench) or narrow bioinformatics tasks. Presently, there remains a lack of benchmarks evaluating agent capability in multi-step data analysis workflows or in solving the machine learning (ML) challenges central to AI-driven therapeutics development, such as perturbation response modeling or drug toxicity prediction. We introduce BioML-bench, the first benchmarking suite for evaluating AI agents on end-to-end biomedical ML tasks.

    Passage read from www.biorxiv.org, which may be a preprint rather than the published version.