On October 7, 2026, Biohub — the philanthropic initiative of Meta CEO Mark Zuckerberg and his wife, Dr. Priscilla Chan — announced that the US government, Meta Platforms and Alphabet (alongside Google DeepMind) had joined its project to build open datasets for training artificial intelligence (AI) models for biological research. According to Reuters, the new commitments bring total investment in the initiative to $1.8 billion. The funding will support the Virtual Biology Initiative, which aims to measure how cells respond to changes across far more conditions than scientists have studied so far, and to use that data to build predictive models that could substantially shorten drug development timelines (a process that currently takes years).

How the funding breaks down

According to Biohub's official announcement, the new partners' funds are allocated as follows. Meta, Google DeepMind and drug-discovery startup Isomorphic Labs are jointly investing $300 million. The US Department of Energy will provide more than $500 million over five years for laboratory measurement, modeling and computation.

The National Institutes of Health (NIH) will coordinate datasets and repositories built with more than $500 million in earlier federal funding; Biohub will standardize them for AI training. Reuters notes that these commitments follow the $500 million Biohub directed to the project in April 2026. The number of major backers of the Virtual Biology Initiative has thus grown sharply, bringing together the largest players from both the public and private sectors.

International collaboration expands

Biohub's official announcement of October 7 confirms that Biohub, the US Department of Energy, the NIH and new funding partners have announced an expansion of the international effort to generate the data that will enable predictive AI models for fighting disease. The official announcement puts total commitments in the new phase at "nearly $2 billion," while Reuters, citing Biohub data, reports the overall figure as $1.8 billion.

The initiative is described as a cross-sector collaboration: government agencies are combining laboratory infrastructure and computing power, while technology companies contribute AI modeling expertise and funding. Biohub presented it as a major expansion of the international push to generate data for biological research.

Open science and embargoes

The datasets will eventually be released publicly, but the companies funding them will get first access. According to Biohub's head of science, Alex Rives, embargo periods will be set for commercial partners — during this time partner groups work with the data, after which it becomes an open scientific resource.

"With commercial funders we have embargo periods where there's a period of time where the groups can work on the data, and then it becomes available as a public scientific resource," said Alex Rives, Biohub's head of science (Reuters).

Biohub describes this arrangement as a way to draw private money into an open-science project. Parallel work funded by the government will carry no such restrictions. In the next phase, Biohub plans to approach pharmaceutical companies and philanthropies — a step intended to further expand the open-science model.

Goals of the five-year plan

The Virtual Biology Initiative's mission is to measure how cells respond to changes across far more conditions than scientists have studied so far, and to use that data to build predictive models that could substantially shorten drug development timelines.

Data will be gathered using a range of methods: spatial transcriptomics, which maps molecular activity inside intact tissue, as well as screens that record how cells respond to changes in their environment. Much of this data has never before been generated in a coordinated way.

"Biology has been just sort of a clever discovery-based science until this point. We have always held this as a community asset, not just for one group, so that it can build upon itself over time," Priscilla Chan said in an interview (Reuters).

According to Alex Rives, current cell datasets cover hundreds of millions of cells, while accurate predictive models will require billions and eventually trillions of cell data points. He noted that such work would normally take decades, while the partners aim to compress it into five years: the first dataset should be ready in about a year, with accurate predictive models expected within five years. Biohub's official announcement likewise notes that the project is aimed at building the foundational data for AI models designed to predict and treat disease (Biohub's official announcement).