
AI models enable cross-species mapping of cell biology
[post_content]
Disclaimer: This article has been automatically aggregated from
Researchers at Stanford Medicine have developed two new artificial intelligence models of the biological cell. The first, called universal cell embedding, paved the way for a second-generation model called TranscriptFormer, which has been trained on data from 112 million cells representing 12 species, ranging from single-celled yeast to humans.
The AI cell models make it dramatically easier to compare cells among species and offer possibilities for new insights into diseases and new cell treatments.
“We’re trying to open up a whole new way to think about cell biology,” said Stephen Quake, Ph.D., a professor of bioengineering. He is co-lead author, along with Jure Leskovec, Ph.D., professor of computer science, and Theo Karaletsos, Ph.D., of the Chan Zuckerberg Initiative, of two scientific articles about the work, published recently in Nature and Science. “These papers mark the beginning of, we hope, a whole new field and a decade of work,” Quake said.
The genome contains the instructions for life, but these days many biologists focus on gene expression—that is, which genes a particular cell is using.
For example, a beta cell in the pancreas expresses the gene for insulin, plus genes related to storage and release of the hormone. The white blood cells known as B cells produce antibodies to fight disease. A cell in your skin, on the other hand, might express genes to make hair or pigment. These differences in expression can be used to tell cell types apart. Healthy and sick cells also have different gene expression patterns, as do cells from different species.
In recent years, scientists have built various atlases, which are databases that define cells and tissues based on their gene expression patterns. The amount of data on tens of thousands of genes for hundreds of millions of cells—from many species—is staggering.
Too much for human brains
“How do we think about all these genes at once? The human brain can’t,” Quake said. “These models help us do that. The algorithms are ways for us to understand very complex data that humans can’t wrap our minds around.”
TranscriptFormer was trained on 12 species. In addition to humans and yeast, the team used cell atlases for mice, rabbits, chickens, zebrafish, fruit flies, the African clawed frog, the parasite that causes malaria, a sea urchin, a sponge and the roundworm Caenorhabditis elegans.
The training worked on gene expression profiles in a way that is analogous to large language models like ChatGPT. LLMs are trained, in essence, by repeatedly looking at texts with some words missing and adjusting their parameters until they can correctly predict the missing words. TranscriptFormer did the same, but with gene expression values instead of words.
Quake describes the resulting model as a “universal space”—not a physical place, but a mathematical space that can be used to measure how similar or different things are. “All the cells of every organism on Earth can sit in that space, and you can look at their relationship to each other,” he said.
One use for TranscriptFormer is to explore evolutionary relationships between different species. For example, scientists who study sponges have wondered which kind of cell in a sponge is most closely related to a neuron in other animals. The model showed that the cell known as a choanocyte was quite similar to neurons in the roundworm and the frog. This helps scientists understand the evolution of neurons and the role of choanocytes in sponges.
The sponge’s neuroid cell, on the other hand, was named that because it has some neuron-like features—but it turns out to have a gene expression profile more like a frog gland, suggesting that it may instead be involved in digestion.
Another use is with new species: When given gene expression data for a species it wasn’t trained on, TranscriptFormer can identify the cell types. The team also showed that it can tell the difference between healthy and diseased cells.
Someday, the model could help guide the design of new treatments that use cells, for example, because it can imagine functional cells that don’t exist right now.
“We hope it’s going to be a very powerful tool for discovery,” Quake said. “This is a tool to help skilled biologists think about complex problems.”
Researchers from the Chan Zuckerberg Initiative and BioHub contributed to the work. The co-first authors of the Nature paper were Yanay Rosen, a doctoral student in computer science at Stanford University, and Yusuf Roohani, Ph.D., now an associate director of machine learning at the Arc Institute.
Publication details
Yanay Rosen et al, Universal cell embedding provides a foundation model for cell biology, Nature (2026). DOI: 10.1038/s41586-026-10689-z
James D. Pearce et al, TranscriptFormer: A generative cell atlas across 1.5 billion years of evolution, Science (2026). DOI: 10.1126/science.aec8514
Provided by
Stanford University Medical Center
Citation:
AI models enable cross-species mapping of cell biology (2026, September 14)
retrieved 14 September 2026
from https://phys.org/news/2026-09-ai-enable-species-cell-biology.html
This document is subject to copyright. Apart from any fair dealing for the purpose of private study or research, no
part may be reproduced without the written permission. The content is provided for information purposes only.
for informational purposes only. We do not claim ownership, accuracy, or liability for the content provided. All rights belong to the original publisher.
