The biggest bet on AI for drug discovery just got a lot bigger. Biohub, the nonprofit research outfit backed by Mark Zuckerberg and Priscilla Chan, said Wednesday that its Virtual Biology Initiative has grown into a $1.8 billion effort, with Google DeepMind, Meta and two US federal agencies now signed on. The goal: software that predicts how a human cell reacts to a drug, a mutation or a disease before anyone runs the experiment.

That’s not a small claim. Drug development today is a decade-long, billion-dollar grind, and most of it is trial and error. The Virtual Biology Initiative wants to replace a chunk of that error with prediction — a flight simulator for medicine, where researchers introduce a variable and the model shows what happens, so only the most promising leads ever reach a petri dish.

Where the $1.8 billion comes from

Biohub put the first $500 million on the table back in April. Wednesday’s announcement roughly triples the project. The new money splits three ways.

Meta, Google DeepMind and the drug-discovery startup Isomorphic Labs are jointly investing $300 million. The Department of Energy is committing more than $500 million over five years in laboratory measurement, modeling and computation. And the National Institutes of Health will coordinate datasets and repositories built with more than $500 million in earlier federal funding, which Biohub will standardize for AI training, according to Reuters.

The corporate lineup is notable. Isomorphic Labs is DeepMind’s drug-discovery spinout, and pairing it with Meta’s compute and the Energy Department’s lab infrastructure gives the initiative coverage across the three things the project needs most: models, measurement and machines.

Why AI drug discovery is starving for data

Modern AI runs on data, and biology has a data problem. Current cell datasets run to hundreds of millions of cells, according to Biohub’s head of science Alex Rives. An accurate predictive model will eventually need billions — and then trillions.

That’s the gap the money is meant to close. The datasets will come from techniques like spatial transcriptomics, which maps molecular activity inside intact tissue, and screens that record how cells respond to changes in their environment. Much of that data has never been generated in a coordinated way at this scale.

“We need to capture the language of biology, we need to capture the language of the cell. And that doesn’t exist today,” Rives told Reuters.

It’s a familiar move from the AI playbook: take a domain where progress has been discovery-based and messy, flood it with standardized measurements, and see whether prediction emerges at scale. Rives said the work would normally take decades and that the partners aim to compress it into five years, with a first dataset ready in about a year.

The labs aren’t the only ones converging on biology. Anthropic has built a wet lab for AI-driven drug development, and the OpenAI Foundation started a grant program worth more than $125 million for biological and medical datasets, Reuters reported. Everybody in frontier AI has decided that biology is next. Nobody has put down $1.8 billion to prove it — until now.

The catch: “open science” with a head start

Biohub calls this an open scientific resource. It mostly is — with an asterisk. The companies writing the checks get embargo periods before the data goes public, a one-year head start to work with the datasets they funded, according to the group’s disclosures. The government-funded work, running in parallel, will carry no such restrictions.

This is the standard tension of public-private science at this scale. The commercial funders need a reason to show up; the public needs to actually benefit. A one-year embargo is Biohub’s answer, and it’s the price of turning a $500 million science bet into a $1.8 billion coalition. Watch whether it holds. The pharmaceutical companies and philanthropies Biohub says it plans to approach next will want their own terms, and each carve-out makes the “community asset” framing harder to sustain.

Why this matters

The honest read is that this is the most serious attempt yet to do for biology what scaled-up data collection did for language — except cells are far less forgiving than sentences. A language model can hallucinate a fact and nobody dies. A cell model that hallucinates a drug response could mislead an entire research program. The predictive models Biohub is aiming for, which Rives expects within five years, will live or die on how well they handle the long tail of cellular behavior that today’s datasets barely cover.

Still, the structure of the bet is right. Biology’s data problem is exactly the kind of problem money can solve: buy the instruments, run the measurements, standardize the results. And the coalition’s composition tells you something about where AI is headed — the frontier isn’t just bigger language models anymore. It’s models of things: proteins, materials, cells. The lab that builds the first working virtual cell won’t just sell drugs faster. It will own the foundation everything downstream is built on.

Sources: Reuters, The Wall Street Journal, THE DECODER.

Frequently asked questions

What is the Virtual Biology Initiative? Biohub’s five-year effort to generate the datasets needed to train AI models that predict how human cells behave under different conditions, with the aim of compressing drug development timelines.

Who is funding the $1.8 billion AI biology effort? Biohub committed $500 million in April 2026. Meta, Google DeepMind and Isomorphic Labs are jointly adding $300 million, the Department of Energy is investing more than $500 million over five years, and NIH is coordinating datasets built with over $500 million in earlier federal funding.

Will the AI biology data be publicly available? Eventually, but commercial funders get an embargo period of about a year before the data they paid for becomes a public scientific resource. Government-funded work will carry no such restrictions.

How could AI cell models speed up drug discovery? Instead of testing every drug candidate in the lab, researchers would run digital experiments on a predictive model of the human cell and move only the most promising leads to physical testing.