Claim 69 · ai-for-science · connection

Evaluation is shifting from question answering to end-to-end tasks: Laurent et al. 2024 caution that high LAB-Bench scores are necessary but not sufficient for useful research assistants, and Miller et al. 2025 introduce BioML-bench precisely because prior agent evaluation was restricted to QA or narrow bioinformatics tasks.

Supported

E17 states human or superhuman LAB-Bench performance is likely necessary but not sufficient for agents to be useful research assistants, and E43 states prior biomedical agent evaluation was largely restricted to question answering or narrow bioinformatics tasks, motivating BioML-bench for end-to-end tasks.

Written by Kimi K3 via Ollama Cloud · checked by GLM-5.3 via Ollama Cloud · 1 Oct, 04:32

Source chain

A

Every quote below was checked, without a model, to appear verbatim in its source.

  1. 01

    “human or superhuman performance on these benchmarks is likely a necessary but not sufficient condition for language models and agents to be useful as assistants for scientific research”

    Full-text passage · p. 2 · evidence E17

    To address the lack of practical task evaluation and provide a benchmark dataset for development of AI systems for scientific research in biology, we introduce here the Language Agent Biology Benchmark, or LAB-Bench. LAB-Bench comprises over 2,400 multiple choice questions that cover important practical research tasks that are largely universal across biology research like recall and reasoning over literature (LitQA2 and SuppQA), interpretation of figures (FigQA) and tables (TableQA), accessing databases (DbQA), writing protocols (ProtocolQA), and comprehending and manipulating DNA and protein sequences (SeqQA, CloningScenarios) (see Table 1 for a breakdown of categories.) The tasks evaluated here therefore represent important early capabilities to benchmark for guiding the future development of useful AI science systems, and we believe that human or superhuman performance on these benchmarks is likely a necessary but not sufficient condition for language models and agents to be useful as assistants for scientific research. 2

    Passage read from arxiv.org, which may be a preprint rather than the published version.

  2. 02

    “their evaluation has largely been restricted to question answering (e.g., LAB-Bench) or narrow bioinformatics tasks”

    Full-text passage · p. 1 · evidence E43

    BioML-bench: Evaluation of AI Agents for End-to-End Biomedical ML Henry E. Miller∗† Shift Bioscience Toronto, Canada Matthew Greenig∗ University of Cambridge ScienceMachine Cambridge, UK Benjamin Tenmann ScienceMachine London, UK Bo Wang University of Toronto Vector Institute Toronto, Canada Abstract Large language model (LLM) agents hold promise for accelerating biomedical research and development (R&D). Several biomedical agents have recently been proposed, but their evaluation has largely been restricted to question answering (e.g., LAB-Bench) or narrow bioinformatics tasks. Presently, there remains a lack of benchmarks evaluating agent capability in multi-step data analysis workflows or in solving the machine learning (ML) challenges central to AI-driven therapeutics development, such as perturbation response modeling or drug toxicity prediction. We introduce BioML-bench, the first benchmarking suite for evaluating AI agents on end-to-end biomedical ML tasks.

    Passage read from www.biorxiv.org, which may be a preprint rather than the published version.