LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Jon M. Laurent, Joseph D. Janizek, Michael Ruzo, Michaela M. Hinks, Michael J. Hammerling, Narayanan, Siddharth, Manvitha Ponnapati, White, Andrew D. and 1 more
Why it has this license class
AChecked 30 Sept 2026. Open license (CC-BY, CC-BY-SA, CC0, public domain): full text indexed and used in synthesis.
| Source | License | Open-access status | Read as |
|---|---|---|---|
| openalex | cc-by-sa | green | Green |
| arxiv | http://creativecommons.org/licenses/by-sa/4.0/ | — | Green |
Abstract
BThere is widespread optimism that frontier Large Language Models (LLMs) and LLM-augmented systems have the potential to rapidly accelerate scientific discovery across disciplines. Today, many benchmarks exist to measure LLM knowledge and reasoning on textbook-style science questions, but few if any benchmarks are designed to evaluate language model performance on practical tasks required for scientific research, such as literature search, protocol planning, and data analysis. As a step toward building such benchmarks, we introduce the Language Agent Biology Benchmark (LAB-Bench), a broad dataset of over 2,400 multiple choice questions for evaluating AI systems on a range of practical biology research capabilities, including recall and reasoning over literature, interpretation of figures, access and navigation of databases, and comprehension and manipulation of DNA and protein sequences. Importantly, in contrast to previous scientific benchmarks, we expect that an AI system that can achieve consistently high scores on the more difficult LAB-Bench tasks would serve as a useful assistant for researchers in areas such as literature search and molecular cloning. As an initial assessment of the emergent scientific task capabilities of frontier language models, we measure performance of several against our benchmark and report results compared to human expert biology researchers. We will continue to update and expand LAB-Bench over time, and expect it to serve as a useful tool in the development of automated research systems going forward. A public subset of LAB-Bench is available for use at the following URL: https://huggingface.co/datasets/futurehouse/lab-bench
Claims built on this paper
D- Evaluation infrastructure is being built alongside the agents themselves, with explicit human comparison as the yardstick: Skarlinski et al. 2024 report PaperQA2 matches or exceeds subject-matter experts on realistic literature tasks, and Laurent et al. 2024's LAB-Bench contributes over 2,400 biology research questions benchmarked against PhD-level scientists. Miller et al. 2025 identify the remaining gap—end-to-end biomedical ML workflows—and introduce BioML-bench to cover it. Trace →
- Evaluation is shifting from question answering to end-to-end tasks: Laurent et al. 2024 caution that high LAB-Bench scores are necessary but not sufficient for useful research assistants, and Miller et al. 2025 introduce BioML-bench precisely because prior agent evaluation was restricted to QA or narrow bioinformatics tasks. Trace →