Core bundle

AI for science and agentic research

Agents that plan, run and write up research, and where they still fall short of experts.

Synthesised claims

A

Each claim cites its evidence; a second model has checked it against that evidence.

  1. 01

    Gottweis et al. 2026's Co-Scientist is a Gemini-based multi-agent system in which agents continuously generate, critique and refine hypotheses under scaled test-time compute, using self-play debate, tournaments and evolution; it was validated on drug repurposing for cancer, liver fibrosis target discovery, and antimicrobial resistance mechanisms.

    Supported2 sources · 1 papersTrace →
  2. 02

    Lu et al. 2026 report that The AI Scientist autonomously performs ideation, experiments, analysis, writing and peer review, and that one generated manuscript passed the first round of peer review at a workshop of a top-tier machine learning conference — though the workshop had a 70% acceptance rate.

    Supported2 sources · 1 papersTrace →
  3. 03

    AI capability is uneven across scientific task types: Skarlinski et al. 2024 show PaperQA2 matches or exceeds experts on literature research tasks, yet Alampara et al. 2025 find vision language models fail at spatial reasoning and cross-modal synthesis, and Minasny et al. 2026 report LLMs answer only up to 65% of advanced soil science exam questions correctly.

    Supported3 sources · 3 papersTrace →
  4. 04

    AI agents have already produced experimentally validated results in multiple domains: Boiko et al. 2023's Coscientist optimized palladium-catalysed cross-couplings, Ghareeb et al. 2026's Robin identified and confirmed ripasudil for dry age-related macular degeneration in vitro, and Swanson et al. 2024's Virtual Lab produced nanobodies with validated binding across SARS-CoV-2 variants.

    Supported3 sources · 3 papersTrace →
  5. 05

    Role-specialized multi-agent teams are a convergent design pattern across systems such as Co-Scientist and the Virtual Lab, and Qi et al. 2026's survey articulates the rationale: planners, executors, validators and critics provide structured redundancy and error-checking that may mitigate hallucination and bias.

    Supported3 sources · 3 papersTrace →
  6. 06

    Evaluation is shifting from question answering to end-to-end tasks: Laurent et al. 2024 caution that high LAB-Bench scores are necessary but not sufficient for useful research assistants, and Miller et al. 2025 introduce BioML-bench precisely because prior agent evaluation was restricted to QA or narrow bioinformatics tasks.

    Supported2 sources · 2 papersTrace →
  7. 07

    The co-scientist paradigm is diffusing beyond biomedicine and chemistry: Minasny et al. 2026 explicitly import the 'AI scientist'/'co-scientist' concept into soil science, and Jamali et al. 2026 envision electron microscopes becoming thinking systems that refine protocols and generate hypotheses.

    Supported2 sources · 2 papersTrace →
  8. 08

    Authors across the bundle flag risks of research automation: Lu et al. 2026 warn of taxing overwhelmed review systems and adding noise to the literature, Li et al. 2025 highlight hallucinations that appear valid but are false, and Resnik et al. 2026 enumerate ethical issues including increasing rates of biased, erroneous and deceptive research.

    Supported3 sources · 3 papersTrace →
  9. 09

    Two complementary strategies turn published literature into agent capability: Huang et al. 2025's Biomni mines tools, databases and protocols from tens of thousands of papers across 25 biomedical domains, while Miao et al. 2026's Paper2Agent converts individual papers into agents that function as virtual corresponding authors.

    Partly supported1 sources · 1 papersTrace →
  10. 10

    Li et al. 2025 explicitly periodize the field, naming Boiko et al. 2023's Coscientist and Huang et al. 2025's Biomni as markers of a 'scientific-agent phase' beginning in 2023 — framing these independently developed systems as milestones of a single transition to autonomous experiment design and workflow iteration.

    Partly supported2 sources · 2 papersTrace →
  11. 11

    Independent teams have converged on multi-agent, tool-using LLM architectures as the core design pattern for AI-for-science systems: Coscientist (Boiko et al. 2023) pairs GPT-4 with search, code execution and lab automation, Co-Scientist (Gottweis et al. 2026) is a multi-agent system built on Gemini, and the Virtual Lab (Swanson et al. 2024) has an LLM principal investigator guiding specialist agents. Qi et al. 2026's survey argues this convergence is functional, since multi-agent redundancy offers error-checking that can mitigate hallucination and bias.

    Supported4 sources · 4 papersTrace →
  12. 12

    The field has shifted from automating narrow, isolated tasks toward general-purpose systems that span much of the research life cycle. Lu et al. 2026's AI Scientist runs from ideation to peer review, Gottweis et al. 2026 frame Co-Scientist as a general collaborator for scientists, and Huang et al. 2025 present Biomni as a general-purpose biomedical agent rather than a specialist workflow.

    Supported3 sources · 3 papersTrace →
  13. 13

    Lu et al. 2026 report that a manuscript generated by The AI Scientist passed the first round of peer review at a workshop of a top-tier machine learning conference, a headline result for end-to-end automation. The same passage notes the workshop had a 70% acceptance rate, which tempers how strong a quality signal this provides.

    Partly supported2 sources · 1 papersTrace →
  14. 14

    Gottweis et al. 2026 show that scaling test-time compute—through self-play debate, hypothesis tournaments and an evolution process—continues to improve hypothesis quality over time. They validate Co-Scientist in three biomedical settings of varied complexity: cancer drug repurposing, liver-fibrosis target discovery, and antimicrobial-resistance mechanism identification.

    Supported2 sources · 1 papersTrace →
  15. 15

    Agent-driven pipelines have already produced experimentally validated wet-lab results in more than one domain: Boiko et al. 2023's Coscientist successfully optimized palladium-catalysed cross-coupling reactions, and Swanson et al. 2024's Virtual Lab designed 92 nanobodies, two of which showed improved binding to recent SARS-CoV-2 variants while retaining binding to the ancestral spike.

    Supported2 sources · 2 papersTrace →
  16. 16

    Evaluation infrastructure is being built alongside the agents themselves, with explicit human comparison as the yardstick: Skarlinski et al. 2024 report PaperQA2 matches or exceeds subject-matter experts on realistic literature tasks, and Laurent et al. 2024's LAB-Bench contributes over 2,400 biology research questions benchmarked against PhD-level scientists. Miller et al. 2025 identify the remaining gap—end-to-end biomedical ML workflows—and introduce BioML-bench to cover it.

    Partly supported3 sources · 3 papersTrace →
  17. 17

    Despite the autonomy claims, current models show sharp reasoning limits that keep humans in the loop: Alampara et al. 2025 find vision-language models fail at spatial reasoning, cross-modal synthesis and multi-step inference, concluding they cannot yet serve as autonomous scientific reasoners. Consistent with this, Minasny et al. 2026 report that LLMs answered only up to 65% of advanced soil-science examination questions.

    Supported3 sources · 2 papersTrace →
  18. 18

    Both system builders and ethicists flag risks to the scientific literature itself: Lu et al. 2026 warn their own technology could tax overwhelmed review systems and add noise to the literature. Resnik et al. 2026 catalogue a broader set of ethical issues, including biased or deceptive research, overreliance on AI, and diffusion of responsibility.

    Supported2 sources · 2 papersTrace →
  19. 19

    Two recent systems treat published papers as machine-actionable substrates rather than static text: Huang et al. 2025's Biomni mines tens of thousands of papers to extract tasks, tools and databases into a unified action space, while Miao et al. 2026's Paper2Agent converts individual papers into MCP-based agents validated to reproduce the original results and answer new queries.

    Supported2 sources · 2 papersTrace →
  20. 20

    Hebenstreit et al. 2026 synthesize four existing autonomy frameworks into a biomedical-specific framework that distinguishes cognitive from physical capabilities. Within that framing, they estimate that general-purpose AI agents combined with continuously operating self-driving laboratories could shrink research project timelines from years to months or weeks.

    Supported2 sources · 1 papersTrace →

Anchor papers

B
PaperYearLicense
Accelerating scientific discovery with Co-Scientist
Nature
Hypothesis generation, validated in vitro.
2026Green
conflict
Towards end-to-end automation of AI research
Nature
The whole research cycle automated, including an accepted workshop paper.
2026Green
The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies
Nature
A team of agents with an AI as principal investigator.
2025Red
Biomni: A General-Purpose Biomedical AI Agent
bioRxiv (Cold Spring Harbor Laboratory)· preprint
A general-purpose biomedical agent.
2025Green
Language agents achieve superhuman synthesis of scientific knowledge
arXiv (Cornell University)· preprint
Literature synthesis: closest to what this demo does.
2024Green
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
arXiv (Cornell University)· preprint
The AI Scientist preprint behind the Nature paper.
2024Green
LAB-Bench: Measuring Capabilities of Language Models for Biology Research
arXiv (Cornell University)· preprint
Where models still trail experts.
2024Green
Autonomous chemical research with large language models
Nature
An agent that plans and runs real experiments.
2023Green
Scientific discovery in the age of artificial intelligence
Nature
Overview and framing.
2023Red
conflict

Found through citations

C
PaperYearLicense
From foundation models to autonomous agents in biology
Genomics Communications
discovered: cites 2 anchor(s) (The Virtual Lab of AI agents designs new SARS-CoV-; Biomni: A General-Purpose Biomedical AI Agent); similarity 0.78
2026Green
Twelve tips for teaching research skills in the age of agentic AI: A guide for health professions educators
Medical Teacher
discovered: cites 1 anchor(s) (Towards end-to-end automation of AI research); similarity 0.80
2026Red
Enhancing soil science research with multi-agent artificial intelligence systems
Frontiers in Science
discovered: cites 2 anchor(s) (Towards end-to-end automation of AI research; Accelerating scientific discovery with Co-Scientis); similarity 0.77
2026Green
Autonomous biomedical research with an artificial intelligence agent
Science
discovered: cites 1 anchor(s) (Accelerating scientific discovery with Co-Scientis); similarity 0.82
2026Red
Autonomous artificial intelligence, scientific research, and human values
AI and Ethics
discovered: cites 2 anchor(s) (The AI Scientist: Towards Fully Automated Open-End; Accelerating scientific discovery with Co-Scientis); similarity 0.76
2026Green
The need for verification in artificial intelligence-driven scientific discovery
Philosophical Transactions of the Royal Society A Mathematical Physical and Engineering Sciences
discovered: cites 1 anchor(s) (The Virtual Lab of AI agents designs new SARS-CoV-); similarity 0.79
2026Green
A multi-agent system for automating scientific discovery
Nature
discovered: cites 3 anchor(s) (Accelerating scientific discovery with Co-Scientis; Language agents achieve superhuman synthesis of sc; Towards end-to-end automation of AI research); similarity 0.70
2026Yellow
Thinking microscopes: agentic AI and the future of electron microscopy
npj Computational Materials
discovered: cites 2 anchor(s) (The Virtual Lab of AI agents designs new SARS-CoV-; Biomni: A General-Purpose Biomedical AI Agent); similarity 0.75
2026Yellow
Artificial Intelligence for Discovery in Life Sciences
Bioconjugate Chemistry
discovered: cites 1 anchor(s) (Towards end-to-end automation of AI research); similarity 0.80
2026Green
What are the limits to biomedical research acceleration through general-purpose AI?
Scientific Reports
discovered: cites 4 anchor(s) (The AI Scientist: Towards Fully Automated Open-End; The Virtual Lab of AI agents designs new SARS-CoV-; Accelerating scientific discovery with Co-Scientis); similarity 0.73
2026Green
Artificial Intelligence agents for biological research: a survey
Briefings in Bioinformatics
discovered: cites 1 anchor(s) (Biomni: A General-Purpose Biomedical AI Agent); similarity 0.80
2026Green
Reimagining research papers as interactive and reliable AI agents
Nature
discovered: cites 3 anchor(s) (Towards end-to-end automation of AI research; The Virtual Lab of AI agents designs new SARS-CoV-; Accelerating scientific discovery with Co-Scientis); similarity 0.75
2026Yellow
Agentic AI and the rise of in silico team science in biomedical research
Nature Biotechnology
discovered: cites 5 anchor(s) (The AI Scientist: Towards Fully Automated Open-End; The Virtual Lab of AI agents designs new SARS-CoV-; LAB-Bench: Measuring Capabilities of Language Mode); similarity 0.65
2026Red
Reimagining biomedical science workflows in the age of large language models
npj Dementia
discovered: cites 2 anchor(s) (Towards end-to-end automation of AI research; Accelerating scientific discovery with Co-Scientis); similarity 0.79
2026Green
Advancing materials discovery through artificial intelligence
Applied Materials Today
discovered: cites 1 anchor(s) (Autonomous chemical research with large language m); similarity 0.79
2025Green
Probing the limitations of multimodal language models for chemistry and materials research
Nature Computational Science
discovered: cites 2 anchor(s) (LAB-Bench: Measuring Capabilities of Language Mode; Language agents achieve superhuman synthesis of sc); similarity 0.78
2025Green
Exploring the use of AI authors and reviewers at Agents4Science
Nature Biotechnology
discovered: cites 3 anchor(s) (The AI Scientist: Towards Fully Automated Open-End; The Virtual Lab of AI agents designs new SARS-CoV-; Accelerating scientific discovery with Co-Scientis); similarity 0.71
2025Red
BioML-bench: Evaluation of AI Agents for End-to-End Biomedical ML
bioRxiv (Cold Spring Harbor Laboratory)· preprint
discovered: cites 2 anchor(s) (The Virtual Lab of AI agents designs new SARS-CoV-; Biomni: A General-Purpose Biomedical AI Agent); similarity 0.76
2025Green
AI as a catalyst for transforming scientific research: a perspective
AI Agent
discovered: cites 1 anchor(s) (The Virtual Lab of AI agents designs new SARS-CoV-); similarity 0.80
2025Green
The rise and potential opportunities of large language model agents in bioinformatics and biomedicine
Briefings in Bioinformatics
discovered: cites 2 anchor(s) (The Virtual Lab of AI agents designs new SARS-CoV-; Biomni: A General-Purpose Biomedical AI Agent); similarity 0.75
2025Yellow
Toward autonomous discovery: agentic AI and the future of ophthalmic research
Current Opinion in Ophthalmology
discovered: cites 1 anchor(s) (The Virtual Lab of AI agents designs new SARS-CoV-); similarity 0.79
2025Red
Agentic Lab: An Agentic-physical AI system for cell and organoid experimentation and manufacturing
bioRxiv (Cold Spring Harbor Laboratory)· preprint
discovered: cites 2 anchor(s) (The Virtual Lab of AI agents designs new SARS-CoV-; Biomni: A General-Purpose Biomedical AI Agent); similarity 0.74
2025Yellow
Empowering biomedical discovery with AI agents
Cell
discovered: cites 2 anchor(s) (Autonomous chemical research with large language m; Scientific discovery in the age of artificial inte); similarity 0.83
2024Green
The Virtual Lab: AI Agents Design New SARS-CoV-2 Nanobodies with Experimental Validation
bioRxiv (Cold Spring Harbor Laboratory)· preprint
discovered: cites 2 anchor(s) (The AI Scientist: Towards Fully Automated Open-End; LAB-Bench: Measuring Capabilities of Language Mode); similarity 0.77
2024Yellow
A review of large language models and autonomous agents in chemistry
Chemical Science
discovered: cites 3 anchor(s) (LAB-Bench: Measuring Capabilities of Language Mode; Autonomous chemical research with large language m; Language agents achieve superhuman synthesis of sc); similarity 0.77
2024Green