Connect with us

NEWS

Mizzou Review Maps Flow Matching Onto Virtual Cells

A Mizzou-led Nature review maps flow matching onto virtual cells, while Arc’s contest showed hybrid models still beating end-to-end AI.

Published

on

A University of Missouri team and outside collaborators published an 18-page review of flow matching, an AI method that learns how biology moves from one state to another. The journal posted it on April 23, 2026. Mizzou issued its news release on September 9.

The university presented the work as a new course for biomedical discovery. The paper itself is a map of a method class that has been growing since 2022, aimed at a prize many labs already chase: an AI virtual cell.

The Nature Paper Maps a Method Already in Use

The April review in Nature Machine Intelligence is titled “Flow matching for generative modelling in bioinformatics and computational biology.” It occupies volume 8, issue 4, pages 517 to 534. The journal site lists 7,963 accesses, 5 citations, and a 20 Altmetric score.

Alex Morehead, a Hopper Postdoctoral Fellow at Lawrence Berkeley National Laboratory and a Cheng lab alumnus, is first author and a corresponding author. Jianlin “Jack” Cheng, a Curators’ Distinguished Professor and Paul K. and Diane Shumaker Professor in Bioinformatics at Mizzou, is the other corresponding author. Six names share equal-contribution marks: Morehead, Lazar Atanackovic, Akshata Hegde, Yanli Wang, Frimpong Boadu, and Joel Selvaraj.

Hegde, Wang, Boadu, Selvaraj, and Cheng are listed with Electrical Engineering and Computer Science and NextGen Precision Health at Mizzou. Atanackovic is listed with the University of Toronto and the Vector Institute. Alexander Tong is at Mila and Université de Montréal. Aditi Krishnapriyan holds appointments at UC Berkeley and Berkeley Lab.

The abstract frames biology as a mapping problem. A diseased cell should be moved toward a healthy one. A sequence should be moved toward a fold. Those maps are hard to write by hand. Flow matching, the authors write, learns a mapping between arbitrary pairs of high-dimensional data distributions and is therefore a fit for molecular and cell biology. The same abstract says each of those uses contributes toward an AI-based virtual cell.

THE FLOW MATCHING CLOCK

  1. October 6, 2022: Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le post the method on arXiv.
  2. 2023: The paper appears at ICLR as a simulation-free way to train continuous normalizing flows.
  3. 2024: The Mizzou-led review counts 8 key molecular flow matching methods in that year alone.
  4. Early 2025: The same review counts 3 foundational methods for single-cell modeling.
  5. April 23, 2026: Nature Machine Intelligence publishes the 18-page review.
  6. September 9, 2026: Mizzou releases a campus news item on the paper.

That gap between the journal date and the campus release is long enough that the field did not wait. New flow matching papers on protein sequence design, ordered water around structures, and RNA-protein motion were still circulating in late August and early September 2026.

What Flow Matching Learns Between States

Most older computer tools in biology freeze a system at one moment. A structure. A snapshot of gene activity. A single microscopy frame. Cheng put the alternative in plain language for the campus release.

Flow matching helps computers learn how biology changes from one state to another. This gives scientists a powerful new way to study everything from protein folding to cell development and cancer progression.

Jianlin Cheng, Curators’ Distinguished Professor, University of Missouri

Lipman and colleagues introduced simulation-free training for continuous flows by teaching a network to match vector fields along fixed probability paths. Those paths can be the same noise-to-data routes diffusion models use. They can also be other routes, including optimal-transport interpolations that run closer to a straight line. On ImageNet, the group reported better likelihood and sample quality than diffusion baselines, with off-the-shelf ODE solvers doing the sampling.

The biology payoff is the free choice of endpoints. The start does not have to be Gaussian noise. It can be an unbound protein, an untreated cell population, or a sequence alphabet. The end can be a bound complex, a perturbed cell state, or a designed molecule. The review points readers to Meta’s 2024 flow matching guide and to Alexander Tong’s PyTorch library for conditional flow matching.

Cheng, a NextGen Precision Health investigator, also said computers can see connections across enormous amounts of data that humans cannot, which helps researchers move faster and ask better questions. That is the sales pitch for any large model. Flow matching’s specific claim is narrower: it learns the path, not only the destination.

Straight Paths, Fewer Solver Steps

Diffusion models, including RFdiffusion for protein backbones, still dominate wet-lab design stories. The review does not throw those models out. It argues flow matching often does the same job with a cleaner training target and a shorter sampling path, especially when the data live on a geometric manifold such as protein frames.

WHERE THE REVIEW SAYS FLOW MATCHING WINS

  • Fewer steps: Sampling follows a learned vector field along a short path, so inference needs fewer network evaluations than a long reverse-noise schedule.
  • Simpler training: The objective regresses a vector field and skips the simulation loop that older continuous-flow models required at train time.
  • Chosen couplings: Researchers can pick how source and target samples are paired, which matters when the start is a real biological state rather than noise.
  • Geometry: Networks can keep SE(3) equivariance and other symmetries that 3D molecules demand, instead of forcing Euclidean noise onto a curved space.

The review walks those ideas through small molecules, proteins, DNA and RNA, their complexes, single-cell phenotypes, and imaging. It treats that stack as scaffolding for a virtual cell, a digital stand-in that would let a lab try an idea on a computer before a bench experiment. Cheng told Mizzou that, over time, this could reduce reliance on animal and human studies and speed progress toward more tailored medicine.

That last claim is the load-bearing one, and it is also the one the field has already put on a public scoreboard.

Arc’s Contest Put Virtual Cells on a Scoreboard

Arc Institute spent 2025 running a Virtual Cell Challenge that asked models to predict how gene activity shifts after a perturbation. Registration opened on June 26, 2025. Winners were named on December 6, 2025, at NeurIPS. More than 5,000 people registered from 114 countries, more than 1,200 teams submitted, and more than 300 teams reached the final round.

First place paid $100,000. It did not go to an end-to-end generative net. Team BM_xTVC from BioMap Research won with xTrimoSCPerturb, a hybrid of deep learning and classical statistics. Arc’s own recap is blunt: pure AI approaches did not beat statistical baselines, so the winners mixed both. Second place paid $50,000 to Sichuan University’s XLearning Lab. Third place paid $25,000 to team Outlier, a Chicago-Dartmouth-Hong Kong group whose TransPert model leaned on summary-level statistics.

VIRTUAL CELL CHALLENGE 2025 PRIZES

Prize Team Model What it actually was
$100,000 first BM_xTVC, BioMap Research xTrimoSCPerturb Hybrid deep learning plus classical statistics
$50,000 second XLearning Lab, Sichuan University Metric-driven conditional generation Pseudo-bulk residuals, not a full cell simulator
$25,000 third Outlier (Chicago, Dartmouth, HKU) TransPert Cross-cell-line stats with similarity-aware aggregation
$100,000 generalist Altos Labs go-with-the-flow Flow matching in gene-expression space

The generalist prize is the detail that cuts both ways for Cheng’s argument. Altos Labs took the generalist prize with go-with-the-flow, a flow matching model that learns time-dependent dynamics directly in gene-expression space, after ranking highest on average across seven metrics. Flow matching can compete on a virtual-cell bench. It did not win the main purse. The main purse went to a team that refused to trust a neural net alone.

Well-funded groups are still building the thing the review treats as a horizon. Biohub, the Chan Zuckerberg biomedical group, listed a unified AI model of the cell among four grand challenges in November 2025 and said it would expand compute tenfold to 10,000 GPUs by 2028. Arc’s STATE model was trained on observational data from 167 million cells and perturbational data from more than 100 million cells across 70 human cell contexts. Those are data and hardware problems as much as modeling problems.

FlowDock Posted a CASP16 Docking Result

Cheng’s group is not a late spectator. His Bioinformatics and Machine Learning Lab has been a regular in CASP, the biennial protein-structure contest, from CASP7 through CASP16, covering 2006 to 2024. In 2012 he was the first to show, in that contest, that deep learning was the best available method for structure prediction. In CASP16 the lab’s MULTICOM predictors ranked No. 1 in Phase 0 protein-complex prediction without stoichiometry, No. 3 in Phase 1 with stoichiometry, No. 2 in tertiary structure, and among the top groups for model-accuracy estimates.

Morehead and Cheng’s FlowDock paper, presented at ISMB 2025 and published in Bioinformatics, is the lab’s own flow matching artifact. It learns to map unbound protein structures to bound complexes for an arbitrary number of ligands, and it emits a confidence score and an affinity estimate with each pose. On the PoseBusters benchmark, FlowDock reports a 51% blind docking success rate from unbound structures and without multiple sequence alignments, ahead of single-sequence AlphaFold 3 on that test. In CASP16’s ligand category it ranked among the top 5 methods for binding-affinity estimates across 140 protein-ligand complexes. Code sits on the lab’s GitHub under BioinfoMachineLearning/FlowDock.

That is the difference between writing a review and shipping a model. The Nature paper surveys FoldFlow, FrameFlow, SemlaFlow, Dirichlet flow matching, AlphaFlow, CellFlow, CryoFM, and a long tail of cousins. FlowDock is the Cheng lab entry on that list. The authors also maintain an open catalog of flow matching methods that they treat as a living companion to the article.

Cheng called flow matching a unifying framework for generative AI in biology, with the potential to change how living systems are modeled and studied. The review’s own language is close to that. A separate University of Illinois survey of the same method class, posted on arXiv, counted more than 30 flow matching papers at NeurIPS 2025 and more than 150 related submissions at ICLR 2026. Unifying, in this case, means crowded.

A Catalog Will Not Buy 10,000 GPUs

The second-order effect of an 18-page Nature review is not a new algorithm. It is a shared syllabus. A graduate student in Columbia, Missouri, can now point to a journal map, a Meta tutorial, Tong’s library, and a GitHub reading list, then try to move cells or backbones along a learned path. That lowers the cost of entry. It does not close the gap with groups that generate their own perturbation atlases and book GPU clusters in the thousands.

WHAT A VIRTUAL CELL STILL REQUIRES

  • Paired states: Flow matching needs a source and a target. Public snapshots are plentiful. Carefully paired healthy-to-diseased or apo-to-holo sets are not.
  • Joint goals: A valid fold is the easy objective. Activity, immune silence, and a process that can be manufactured pull in different directions, and the training data for that joint map is thin.
  • Bench truth: Arc’s 2025 contest showed that a naive statistical baseline still disciplines pure nets on perturbation prediction.
  • Compute: Biohub’s published plan is 10,000 GPUs by 2028, a scale no review paper can substitute for.

The harder product is not one more generator of plausible backbones. It is a model that can be pointed at two biological states and trusted on the path between them when the objective is a drug that also has to express, stay quiet in an immune system, and come off a production line. Flow matching is a candidate because the start and end distributions can be chosen. Whether those joint datasets exist yet is a separate question, and Arc’s leaderboard suggests they mostly do not.

Cheng’s long-term virtual cell would let a lab test ideas on a computer before animals or patients. The review is honest about that horizon. The contest results are honest in the other direction. Hybrid systems that keep classical statistics in the loop still won the cash prize built to test this exact dream.

The GitHub catalog will need another edition. On September 8, 2026, a Dirichlet flow matching method for structure-conditioned sequence design was already making the rounds, with wet-lab nanobody tests attached. The map Mizzou is now promoting is a good map of a field that has not stopped moving.

Harry is the editor and publisher of MY WORLD NEWS 24, an independent title under his own ownership. Ten years of reporting and then editing taught him that a global readership is not served by assuming everyone lives in the same country. Stories here state currencies, units and time zones explicitly, name the country a law or a company belongs to, and explain local context rather than treating it as known. That care extends to sourcing: a claim is anchored to the filing, statement, transcript or dataset that made it, wherever in the world it was issued, and each figure is checked against that source before publication. The site reports news, business and technology, science and sports, entertainment, lifestyle and travel, and auto and gaming, all with the same standard of evidence. Mistakes are fixed under a corrections policy anyone can read, and the page carries a note saying what was changed. Harry reads every message sent by readers and replies from support@myworldnews24.com.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending