From foundation models to autonomous agents in biology
Shenghui Huang, Mei Lang, Zihan Chen, Chenxu Yang, Xiaoying Huang, Zeynab Mohtashaminia, Yuzhong Peng
Why it has this license class
AChecked 30 Sept 2026. Open license (CC-BY, CC-BY-SA, CC0, public domain): full text indexed and used in synthesis.
| Source | License | Open-access status | Read as |
|---|---|---|---|
| openalex | cc-by | diamond | Green |
| unpaywall | cc-by | gold | Green |
Abstract
BAdvances in sequencing and multi-omics have unleashed exponential biological data growth, from genomes and transcriptomes to single-cell and spatial profiles. Traditional pipelines, reliant on manual curation, strain under this deluge. This bottleneck hampers discovery from terabyte-scale datasets. Large Language Models (LLMs) and AI agents are emerging as a powerful paradigm to address these challenges. Breakthroughs in foundation models pre-trained on biological 'languages' offer in-context learning and generative capabilities far beyond prior bioinformatics tools. By coupling LLM reasoning with multi-agent systems, Retrieval-Augmented Generation (RAG), and the Model Context Protocol (MCP), autonomous AI research agents can plan experiments, execute analyses, and even generate hypotheses with minimal human guidance. While these technologies promise to augment human intellect, their deployment presents critical challenges in reliability, biosecurity, and accessibility. Navigating these obstacles is key to ushering in an era of accelerated discovery and personalized medicine. By moving from static models to active agents, we are witnessing the rise of the 'digital biologist'—an AI collaborator poised to reshape biomedical research. We trace this paradigm's rapid evolution, from foundation models learning the language of biological sequences and single-cell data, to autonomous agents capable of automating analysis, designing experiments, and driving drug discovery. By synthesizing these developments, we offer a strategic roadmap for researchers to navigate the opportunities and challenges of this AI-driven era. Finally, to support the community, a public, actively maintained resource list of models, agents, and datasets is available on our project website: http://awesomebio.webioinfo.top.
Claims built on this paper
DNone yet.