LIVE
Google DeepMind Introduces Gemini 3.8 Live and Its Extended Thinking Variant15/09/26 · Google DeepMind|What's at stake in AI's trillion-dollar infrastructure bet15/09/26|A Flaw in Chain-of-Thought Safety Monitoring14/09/26|Stellar Colosseum: A Multi-Agent System for Long-Horizon Mathematical Research14/09/26|Apple Code Hints Siri Could Be Swapped for ChatGPT or Claude14/09/26 · Apple|Anthropic says Houthi-linked actors used Claude Code for missile guidance software13/09/26 · Anthropic|Yoshua Bengio examines why AI agents lie, cheat and coordinate13/09/26|Terence Tao warns of a 'severe misalignment' in OpenAI's mathematical claims11/09/26 · OpenAI|Terence Tao warns of a 'severe misalignment' of AI in mathematics11/09/26|OpenAI's Navier-Stokes release comes with a Lean 4 formal proof10/09/26 · OpenAI|OpenAI publishes documentation for its Agents API10/09/26 · OpenAI|OpenAI Launches the Agents API, a Managed Service for Cloud Agents10/09/26 · OpenAI|Google DeepMind Introduces Gemini 3.8 Live and Its Extended Thinking Variant15/09/26 · Google DeepMind|What's at stake in AI's trillion-dollar infrastructure bet15/09/26|A Flaw in Chain-of-Thought Safety Monitoring14/09/26|Stellar Colosseum: A Multi-Agent System for Long-Horizon Mathematical Research14/09/26|Apple Code Hints Siri Could Be Swapped for ChatGPT or Claude14/09/26 · Apple|Anthropic says Houthi-linked actors used Claude Code for missile guidance software13/09/26 · Anthropic|Yoshua Bengio examines why AI agents lie, cheat and coordinate13/09/26|Terence Tao warns of a 'severe misalignment' in OpenAI's mathematical claims11/09/26 · OpenAI|Terence Tao warns of a 'severe misalignment' of AI in mathematics11/09/26|OpenAI's Navier-Stokes release comes with a Lean 4 formal proof10/09/26 · OpenAI|OpenAI publishes documentation for its Agents API10/09/26 · OpenAI|OpenAI Launches the Agents API, a Managed Service for Cloud Agents10/09/26 · OpenAI|
BusinessOpenAI

OpenAI is funding the creation of biological data to train its models

OpenAI is paying to generate new biological data, drawing in part on a proposal to recover archives from bankrupt biotech firms to enrich AI model training.

September 15, 20263 min readPublished byMIT Tech Review

The scarcity of high-quality biological data has long constrained efforts to build AI systems capable of accelerating medical research. Unlike text or images, which are abundant online, experimental biology data is expensive to generate, scattered across labs, and frequently shielded by trade secrecy. OpenAI appears to be addressing this bottleneck directly by funding the creation of new datasets rather than relying solely on what already exists.

The MIT Technology Review piece revisits a proposal made last year by clinical trial policy analyst Ruxandra Teslo. Her idea involved tapping into the bankruptcy proceedings of failed biotech companies to obtain detailed regulatory filings, manufacturing strategies, and safety data, information typically kept confidential as valuable trade secrets. By bidding on these archives during liquidation, it could become possible to assemble datasets that would otherwise remain out of reach.

This approach highlights a broader challenge facing AI applied to the life sciences: useful data does exist, but much of it sits locked away in commercial or legal silos. OpenAI's decision to fund the generation of new biological data suggests a shift in strategy, moving beyond aggregating public sources toward actively investing in the production or acquisition of materials previously inaccessible.

While the specifics of OpenAI's program remain unclear at this stage, the initiative fits a broader pattern among major AI labs seeking proprietary data in specialized domains as generic text reserves become depleted. For the biotech sector, this could also spark a debate over how to value and assign ownership to data once treated as worthless byproducts of a company's collapse.

Tags
openaibiologybiotechdatalife-sciencesdrug-discovery

Read also