Biohub, the US Department of Energy, the National Institutes of Health, Google DeepMind, Isomorphic Labs, and Meta committed $1.8 billion on October 7 to generate open biological data to train AI models. Biohub calls it the largest coordinated commitment to AI-ready biological data to date.
The target is a virtual cell: a model precise enough that a researcher could run an experiment digitally, predicting how a cell will respond to a drug or an intervention, then only verify the promising results in a lab.
Where the money comes from
The $1.8 Billions is not a single pot, and the four contributions are different in kind.
| Contributor | Amount | What it covers |
|---|---|---|
| Biohub | $500m | Founding commitment, new measurement technology and external research |
| DO | $500m + over 5 years | Laboratory measurement, modeling, calculation |
| Google DeepMind, Isomorphic Labs, Meta | $300m combined | Technologies and multimodal datasets |
| NIH | $500m+ | Data sets and repositories of past federal investments |
Read the NIH line carefully. The half billion describes money already spent. NIH’s role is to coordinate datasets, repositories and knowledge bases developed through more than $500 million in previous federal investment, which Biohub then standardizes for AI training.
That’s really valuable work, because existing biomedical data is scattered across incompatible formats and standardization is the biggest part of the problem. But it’s not new spending, and the $1.8 billion headline counts it alongside money not yet committed to anything. New funding is closer to $1.3 billion.
Which buys Biohub its own $500 million
The breakdown here is more specific than most announcements provide.
$400 Millions go to technologies that expand what biologists can measure: Cryo-electron tomography that resolves close to atomic details within cells, microscopy capable of imaging millions to billions of cells in living tissue, and engineering tools to disrupt biology at the molecular, cellular, tissue, and whole-organism levels.
The remaining $100 Millions of funds research outside Biohub.
The government half
DO Contribution runs through the Genesis Mission, a cross-agency initiative that it leads. The resources it brings are physical rather than financial: exascale supercomputing, X-ray and neutron scattering, cryo-electron microscopy and tomography, and autonomous laboratories across the National Laboratory System, including the Joint Genome Institute and the Environmental Molecular Sciences Laboratory.
NIH works through its Bio Genesis mission, pulling together national repositories cataloged by the National Library of Medicine and the National Center for Biotechnology Information, plus Common Fund programs that are already building biological atlases and common data standards.
Darío Gil, DOE’s Under Secretary for Science, framed it as combining the department’s computing and measurement assets with Biohub’s AI capability to set a new standard for open science.
Who else is involved
The institutional list is broader than your average press release.
The Allen Institute, Broad Institute, Gladstone Institute, the Human Cell Atlas, the Human Protein Atlas and the Wellcome Sanger Institute joined in organizing the scientific community around the effort. Biohub notes that these groups are experienced in coordinating international collaboration, which goes back to the Human Genome Project.
NVIDIA provides accelerated computing infrastructure, domain-specific software and technical expertise. Renaissance Philanthropy is working to expand funding for data generation.
The hard part is not the money
Biohub is straightforward about what it builds, and it’s less glamorous than the AI framing suggests.
The work is a layer that keeps everyone’s data sets together: common standards, common identifiers and a single access point. Anyone who has tried to combine biological datasets from two institutions will recognize why that is the bottleneck rather than computing or funding.
Biohub was formed here. It built and maintains CELLxGENE and the CryoET Data Portal, and runs projects including Tabula Sapiens, OpenCell and Zebrahub.
Alex Rives, Biohub’s Head of Science, called the virtual cell one of the most important challenges for the next era of science, saying it requires coordinated data generation on a national and international scale. Pushmeet Kohli from Google DeepMind made the same point from the other direction: the challenge will not be solved without open experimental data on an unprecedented scale showing how living cells respond to change.
What to watch
No timeline has been published for when usable models or datasets will arrive, and no technical specification for the data standards.
Those two things will determine if this works. A standardized, open-access biological data commons would be really useful for researchers everywhere, including in India, where access to large-scale experimental data is a real constraint on biomedical work. A commons that exists primarily as a funding announcement would not.
The organizations involved have the money and the tools. Whether they will agree on formats is the question that no one has yet answered.
