AI for science and agentic research
Agents that plan, run and write up research, and where they still fall short of experts.
Synthesised claims
AEach claim cites its evidence; a second model has checked it against that evidence.
- 01
Gottweis et al. 2026's Co-Scientist is a Gemini-based multi-agent system in which agents continuously generate, critique and refine hypotheses under scaled test-time compute, using self-play debate, tournaments and evolution; it was validated on drug repurposing for cancer, liver fibrosis target discovery, and antimicrobial resistance mechanisms.
- 02
Lu et al. 2026 report that The AI Scientist autonomously performs ideation, experiments, analysis, writing and peer review, and that one generated manuscript passed the first round of peer review at a workshop of a top-tier machine learning conference — though the workshop had a 70% acceptance rate.
- 03
AI capability is uneven across scientific task types: Skarlinski et al. 2024 show PaperQA2 matches or exceeds experts on literature research tasks, yet Alampara et al. 2025 find vision language models fail at spatial reasoning and cross-modal synthesis, and Minasny et al. 2026 report LLMs answer only up to 65% of advanced soil science exam questions correctly.
- 04
AI agents have already produced experimentally validated results in multiple domains: Boiko et al. 2023's Coscientist optimized palladium-catalysed cross-couplings, Ghareeb et al. 2026's Robin identified and confirmed ripasudil for dry age-related macular degeneration in vitro, and Swanson et al. 2024's Virtual Lab produced nanobodies with validated binding across SARS-CoV-2 variants.
- 05
Role-specialized multi-agent teams are a convergent design pattern across systems such as Co-Scientist and the Virtual Lab, and Qi et al. 2026's survey articulates the rationale: planners, executors, validators and critics provide structured redundancy and error-checking that may mitigate hallucination and bias.
- 06
Evaluation is shifting from question answering to end-to-end tasks: Laurent et al. 2024 caution that high LAB-Bench scores are necessary but not sufficient for useful research assistants, and Miller et al. 2025 introduce BioML-bench precisely because prior agent evaluation was restricted to QA or narrow bioinformatics tasks.
- 07
The co-scientist paradigm is diffusing beyond biomedicine and chemistry: Minasny et al. 2026 explicitly import the 'AI scientist'/'co-scientist' concept into soil science, and Jamali et al. 2026 envision electron microscopes becoming thinking systems that refine protocols and generate hypotheses.
- 08
Authors across the bundle flag risks of research automation: Lu et al. 2026 warn of taxing overwhelmed review systems and adding noise to the literature, Li et al. 2025 highlight hallucinations that appear valid but are false, and Resnik et al. 2026 enumerate ethical issues including increasing rates of biased, erroneous and deceptive research.
- 09
Two complementary strategies turn published literature into agent capability: Huang et al. 2025's Biomni mines tools, databases and protocols from tens of thousands of papers across 25 biomedical domains, while Miao et al. 2026's Paper2Agent converts individual papers into agents that function as virtual corresponding authors.
- 10
Li et al. 2025 explicitly periodize the field, naming Boiko et al. 2023's Coscientist and Huang et al. 2025's Biomni as markers of a 'scientific-agent phase' beginning in 2023 — framing these independently developed systems as milestones of a single transition to autonomous experiment design and workflow iteration.
- 11
Independent teams have converged on multi-agent, tool-using LLM architectures as the core design pattern for AI-for-science systems: Coscientist (Boiko et al. 2023) pairs GPT-4 with search, code execution and lab automation, Co-Scientist (Gottweis et al. 2026) is a multi-agent system built on Gemini, and the Virtual Lab (Swanson et al. 2024) has an LLM principal investigator guiding specialist agents. Qi et al. 2026's survey argues this convergence is functional, since multi-agent redundancy offers error-checking that can mitigate hallucination and bias.
- 12
The field has shifted from automating narrow, isolated tasks toward general-purpose systems that span much of the research life cycle. Lu et al. 2026's AI Scientist runs from ideation to peer review, Gottweis et al. 2026 frame Co-Scientist as a general collaborator for scientists, and Huang et al. 2025 present Biomni as a general-purpose biomedical agent rather than a specialist workflow.
- 13
Lu et al. 2026 report that a manuscript generated by The AI Scientist passed the first round of peer review at a workshop of a top-tier machine learning conference, a headline result for end-to-end automation. The same passage notes the workshop had a 70% acceptance rate, which tempers how strong a quality signal this provides.
- 14
Gottweis et al. 2026 show that scaling test-time compute—through self-play debate, hypothesis tournaments and an evolution process—continues to improve hypothesis quality over time. They validate Co-Scientist in three biomedical settings of varied complexity: cancer drug repurposing, liver-fibrosis target discovery, and antimicrobial-resistance mechanism identification.
- 15
Agent-driven pipelines have already produced experimentally validated wet-lab results in more than one domain: Boiko et al. 2023's Coscientist successfully optimized palladium-catalysed cross-coupling reactions, and Swanson et al. 2024's Virtual Lab designed 92 nanobodies, two of which showed improved binding to recent SARS-CoV-2 variants while retaining binding to the ancestral spike.
- 16
Evaluation infrastructure is being built alongside the agents themselves, with explicit human comparison as the yardstick: Skarlinski et al. 2024 report PaperQA2 matches or exceeds subject-matter experts on realistic literature tasks, and Laurent et al. 2024's LAB-Bench contributes over 2,400 biology research questions benchmarked against PhD-level scientists. Miller et al. 2025 identify the remaining gap—end-to-end biomedical ML workflows—and introduce BioML-bench to cover it.
- 17
Despite the autonomy claims, current models show sharp reasoning limits that keep humans in the loop: Alampara et al. 2025 find vision-language models fail at spatial reasoning, cross-modal synthesis and multi-step inference, concluding they cannot yet serve as autonomous scientific reasoners. Consistent with this, Minasny et al. 2026 report that LLMs answered only up to 65% of advanced soil-science examination questions.
- 18
Both system builders and ethicists flag risks to the scientific literature itself: Lu et al. 2026 warn their own technology could tax overwhelmed review systems and add noise to the literature. Resnik et al. 2026 catalogue a broader set of ethical issues, including biased or deceptive research, overreliance on AI, and diffusion of responsibility.
- 19
Two recent systems treat published papers as machine-actionable substrates rather than static text: Huang et al. 2025's Biomni mines tens of thousands of papers to extract tasks, tools and databases into a unified action space, while Miao et al. 2026's Paper2Agent converts individual papers into MCP-based agents validated to reproduce the original results and answer new queries.
- 20
Hebenstreit et al. 2026 synthesize four existing autonomy frameworks into a biomedical-specific framework that distinguishes cognitive from physical capabilities. Within that framing, they estimate that general-purpose AI agents combined with continuously operating self-driving laboratories could shrink research project timelines from years to months or weeks.
Anchor papers
B| Paper | Year | License |
|---|---|---|
| Accelerating scientific discovery with Co-Scientist Nature Hypothesis generation, validated in vitro. | 2026 | Green conflict |
| Towards end-to-end automation of AI research Nature The whole research cycle automated, including an accepted workshop paper. | 2026 | Green |
| The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies Nature A team of agents with an AI as principal investigator. | 2025 | Red |
| Biomni: A General-Purpose Biomedical AI Agent bioRxiv (Cold Spring Harbor Laboratory)· preprint A general-purpose biomedical agent. | 2025 | Green |
| Language agents achieve superhuman synthesis of scientific knowledge arXiv (Cornell University)· preprint Literature synthesis: closest to what this demo does. | 2024 | Green |
| The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery arXiv (Cornell University)· preprint The AI Scientist preprint behind the Nature paper. | 2024 | Green |
| LAB-Bench: Measuring Capabilities of Language Models for Biology Research arXiv (Cornell University)· preprint Where models still trail experts. | 2024 | Green |
| Autonomous chemical research with large language models Nature An agent that plans and runs real experiments. | 2023 | Green |
| Scientific discovery in the age of artificial intelligence Nature Overview and framing. | 2023 | Red conflict |
Found through citations
C| Paper | Year | License |
|---|---|---|
| From foundation models to autonomous agents in biology Genomics Communications discovered: cites 2 anchor(s) (The Virtual Lab of AI agents designs new SARS-CoV-; Biomni: A General-Purpose Biomedical AI Agent); similarity 0.78 | 2026 | Green |
| Twelve tips for teaching research skills in the age of agentic AI: A guide for health professions educators Medical Teacher discovered: cites 1 anchor(s) (Towards end-to-end automation of AI research); similarity 0.80 | 2026 | Red |
| Enhancing soil science research with multi-agent artificial intelligence systems Frontiers in Science discovered: cites 2 anchor(s) (Towards end-to-end automation of AI research; Accelerating scientific discovery with Co-Scientis); similarity 0.77 | 2026 | Green |
| Autonomous biomedical research with an artificial intelligence agent Science discovered: cites 1 anchor(s) (Accelerating scientific discovery with Co-Scientis); similarity 0.82 | 2026 | Red |
| Autonomous artificial intelligence, scientific research, and human values AI and Ethics discovered: cites 2 anchor(s) (The AI Scientist: Towards Fully Automated Open-End; Accelerating scientific discovery with Co-Scientis); similarity 0.76 | 2026 | Green |
| The need for verification in artificial intelligence-driven scientific discovery Philosophical Transactions of the Royal Society A Mathematical Physical and Engineering Sciences discovered: cites 1 anchor(s) (The Virtual Lab of AI agents designs new SARS-CoV-); similarity 0.79 | 2026 | Green |
| A multi-agent system for automating scientific discovery Nature discovered: cites 3 anchor(s) (Accelerating scientific discovery with Co-Scientis; Language agents achieve superhuman synthesis of sc; Towards end-to-end automation of AI research); similarity 0.70 | 2026 | Yellow |
| Thinking microscopes: agentic AI and the future of electron microscopy npj Computational Materials discovered: cites 2 anchor(s) (The Virtual Lab of AI agents designs new SARS-CoV-; Biomni: A General-Purpose Biomedical AI Agent); similarity 0.75 | 2026 | Yellow |
| Artificial Intelligence for Discovery in Life Sciences Bioconjugate Chemistry discovered: cites 1 anchor(s) (Towards end-to-end automation of AI research); similarity 0.80 | 2026 | Green |
| What are the limits to biomedical research acceleration through general-purpose AI? Scientific Reports discovered: cites 4 anchor(s) (The AI Scientist: Towards Fully Automated Open-End; The Virtual Lab of AI agents designs new SARS-CoV-; Accelerating scientific discovery with Co-Scientis); similarity 0.73 | 2026 | Green |
| Artificial Intelligence agents for biological research: a survey Briefings in Bioinformatics discovered: cites 1 anchor(s) (Biomni: A General-Purpose Biomedical AI Agent); similarity 0.80 | 2026 | Green |
| Reimagining research papers as interactive and reliable AI agents Nature discovered: cites 3 anchor(s) (Towards end-to-end automation of AI research; The Virtual Lab of AI agents designs new SARS-CoV-; Accelerating scientific discovery with Co-Scientis); similarity 0.75 | 2026 | Yellow |
| Agentic AI and the rise of in silico team science in biomedical research Nature Biotechnology discovered: cites 5 anchor(s) (The AI Scientist: Towards Fully Automated Open-End; The Virtual Lab of AI agents designs new SARS-CoV-; LAB-Bench: Measuring Capabilities of Language Mode); similarity 0.65 | 2026 | Red |
| Reimagining biomedical science workflows in the age of large language models npj Dementia discovered: cites 2 anchor(s) (Towards end-to-end automation of AI research; Accelerating scientific discovery with Co-Scientis); similarity 0.79 | 2026 | Green |
| Advancing materials discovery through artificial intelligence Applied Materials Today discovered: cites 1 anchor(s) (Autonomous chemical research with large language m); similarity 0.79 | 2025 | Green |
| Probing the limitations of multimodal language models for chemistry and materials research Nature Computational Science discovered: cites 2 anchor(s) (LAB-Bench: Measuring Capabilities of Language Mode; Language agents achieve superhuman synthesis of sc); similarity 0.78 | 2025 | Green |
| Exploring the use of AI authors and reviewers at Agents4Science Nature Biotechnology discovered: cites 3 anchor(s) (The AI Scientist: Towards Fully Automated Open-End; The Virtual Lab of AI agents designs new SARS-CoV-; Accelerating scientific discovery with Co-Scientis); similarity 0.71 | 2025 | Red |
| BioML-bench: Evaluation of AI Agents for End-to-End Biomedical ML bioRxiv (Cold Spring Harbor Laboratory)· preprint discovered: cites 2 anchor(s) (The Virtual Lab of AI agents designs new SARS-CoV-; Biomni: A General-Purpose Biomedical AI Agent); similarity 0.76 | 2025 | Green |
| AI as a catalyst for transforming scientific research: a perspective AI Agent discovered: cites 1 anchor(s) (The Virtual Lab of AI agents designs new SARS-CoV-); similarity 0.80 | 2025 | Green |
| The rise and potential opportunities of large language model agents in bioinformatics and biomedicine Briefings in Bioinformatics discovered: cites 2 anchor(s) (The Virtual Lab of AI agents designs new SARS-CoV-; Biomni: A General-Purpose Biomedical AI Agent); similarity 0.75 | 2025 | Yellow |
| Toward autonomous discovery: agentic AI and the future of ophthalmic research Current Opinion in Ophthalmology discovered: cites 1 anchor(s) (The Virtual Lab of AI agents designs new SARS-CoV-); similarity 0.79 | 2025 | Red |
| Agentic Lab: An Agentic-physical AI system for cell and organoid experimentation and manufacturing bioRxiv (Cold Spring Harbor Laboratory)· preprint discovered: cites 2 anchor(s) (The Virtual Lab of AI agents designs new SARS-CoV-; Biomni: A General-Purpose Biomedical AI Agent); similarity 0.74 | 2025 | Yellow |
| Empowering biomedical discovery with AI agents Cell discovered: cites 2 anchor(s) (Autonomous chemical research with large language m; Scientific discovery in the age of artificial inte); similarity 0.83 | 2024 | Green |
| The Virtual Lab: AI Agents Design New SARS-CoV-2 Nanobodies with Experimental Validation bioRxiv (Cold Spring Harbor Laboratory)· preprint discovered: cites 2 anchor(s) (The AI Scientist: Towards Fully Automated Open-End; LAB-Bench: Measuring Capabilities of Language Mode); similarity 0.77 | 2024 | Yellow |
| A review of large language models and autonomous agents in chemistry Chemical Science discovered: cites 3 anchor(s) (LAB-Bench: Measuring Capabilities of Language Mode; Autonomous chemical research with large language m; Language agents achieve superhuman synthesis of sc); similarity 0.77 | 2024 | Green |