
Share
With a $250 billion equity stake looming, the OpenAI Foundation is funding efforts to mine bankruptcy filings and clinical trial archives, betting that scarce biological data, not raw model power, is medicine's real bottleneck.
The thesis behind OpenAI's newest philanthropic push is simple: AI models will not cure disease faster unless they are fed far more biological data than currently exists in usable form. That is the premise behind Public Data for Health, a new initiative from the OpenAI Foundation, the nonprofit parent of OpenAI, which announced this week it will pay to create "high-quality scientific datasets" for medical research.
The idea traces back to Ruxandra Teslo, a policy analyst focused on clinical trials, who last year proposed something unconventional. Bid at bankruptcy proceedings of failed biotech companies, she argued, and you can obtain regulatory filings, manufacturing strategies, and safety data normally locked away as trade secrets. She dubbed this trove "biotech's lost archive." The OpenAI Foundation is now funding that idea directly, awarding $500,000 to 1Day Sooner, an advocacy group for clinical trial volunteers, which Teslo advises, to pursue it.
The math behind the opportunity is notable. Josh Morrison, 1Day Sooner's president and cofounder, says nonexclusive copies of bankrupt biotech datasets could be acquired for "a few tens of thousands of dollars" each. His organization already holds three such datasets, two of them donated by Lumen Bioscience, a biotech that previously used Chapter 11 proceedings to gain insight into a rival's drug development work. Two other bids this year did not succeed.
Morgan Levine, a former vice president for computation at longevity company Altos Labs, frames the problem bluntly: "Everyone is recognizing that data is the biggest bottleneck in successfully applying AI to biology." That statement cuts against the industry's dominant narrative, which has largely fixated on model scale and compute as the binding constraints.
Teslo's argument sharpens the point further. Roughly 70% of the money and time in drug development, she says, goes into clinical development: organizing trials and testing compounds. Yet that process remains "basically a black box, especially for small biotech companies generating the innovations." Her wager is that a large enough archive of common technical documents, the files that capture the full regulatory back-and-forth between companies and agencies, could train an AI to function as a genuine regulatory copilot rather than a generic chatbot bolted onto a lab notebook.
The grant to 1Day Sooner is one of several announced under the new program. The OpenAI Foundation is putting $40 million into a University of North Carolina, Chapel Hill effort to collect data on novel cancer vaccines, and it is backing OpenAdmet, a group that runs competitions where researchers try to predict drug effects. None of these figures individually rival OpenAI's largest single gift to date: $100 million awarded in August to the Common Health Coalition, which helps patients access hepatitis C treatments.

Jacob Trefethen, an executive at the foundation, describes the organization's approach as deliberately hands-off. "We're starting grantmaking when we think the best way to achieve that mission is to make grants to external nonprofits, research institutions, and other third parties," he said. The foundation is targeting $1 billion in giving by the end of the year, a pace that would be aggressive for any philanthropic entity, let alone one still building out its leadership team.
The scale of capital involved is what makes this story more than a niche research-funding update. OpenAI's for-profit arm is reportedly preparing for an IPO that could value the company at $1 trillion, according to Reuters. The Foundation holds a 26% equity stake in that entity, which puts it on track to become the wealthiest charitable organization on the planet, potentially sitting on $250 billion in stock value. For context, the Gates Foundation and its associated trust held roughly $180 billion at the end of 2025. If that valuation materializes, the Foundation's grantmaking budget could dwarf established players in global health philanthropy within a few years.
None of this is happening in a vacuum free of risk. Bankruptcy data acquisition is emerging as what some are calling a "new land grab" for AI training material. Google recently won a bid for the corporate data of failed carrier Spirit Airlines, including 100 million emails, a deal that drew objections from flight attendants and privacy advocates concerned about proprietary information being swept up in the sale. The same dynamic could apply to biotech bankruptcies: common technical documents contain detailed scientific and medical measurements that companies would ordinarily guard closely, and the ethics of harvesting them post-failure are not fully settled.
There is also a broader unease shadowing the entire AI-biology intersection. Fears have surfaced, some stoked by AI company insiders themselves, that advanced AI could be used to engineer a bioweapon capable of mass harm. Estimates circulating among researchers put the odds of human extinction within a decade at 10% or more, according to public statements from industry figures. Sam Altman and Elon Musk both endorsed a call last week from Anthropic CEO Dario Amodei to slow the pace of AI capability development so safety measures can catch up. That OpenAI is simultaneously racing to feed models more biological data, while its leadership publicly worries about AI-enabled biological catastrophe, is a tension the industry has not resolved.
The data-scarcity thesis is credible and grounded in specifics: real dollar figures for bankruptcy data acquisition, a defined 70% share of drug development time spent in clinical trials, and named institutional grants totaling well over $140 million so far. But the Foundation's ambitions are still in their earliest phase, with key roles unfilled and only one large gift, the $100 million hepatitis C award, disbursed to date. Investors and observers watching OpenAI's structure should note the unusual setup: a nonprofit sitting on a potential $250 billion equity stake, moving capital into biomedical data infrastructure just as the parent company prepares for a trillion-dollar IPO. Whether that capital translates into genuine breakthroughs, or simply subsidizes a data land grab with uncertain safety guardrails, will depend on execution over the next 12 to 24 months.
Tags
Original Sources
AI models need more data about biology, and OpenAI is paying to create it
↗ https://www.technologyreview.com/2026/09/15/1144129/ai-models-need-more-data-about-biology-and-openai-is-paying-to-create-it
About the author
Marcus began tracking AI's market implications in 2016, noticing AI-related patent filings accelerating ahead of earnings upgrades before most of the sell-side had caught on. A former fixed-income quantitative analyst, he spent two decades building models that priced risk across emerging markets before pivoting to cover the economic impact of AI full-time. His writing translates opaque technical developments into clear risk/reward terms — and he's rarely diplomatic about the gap between AI valuations and underlying fundamentals. He believes most market participants still underestimate AI's long-run deflationary effect on knowledge work.
More from The Analyst →This Week's Edition
16 September 2026
31 articles
Related Articles
Related Articles
More Stories
© 2026 Cedar & Bloom. All rights reserved.