AI工具Score B (52)

AI models need more data about biology, and OpenAI is paying to create it

1 小时前2 viewsSource: MIT Technology Review

Last year Ruxandra Teslo, a policy analyst who focuses on clinical trials, posted an idea for supercharging medical AI systems: Use data from failed biotech companies.

By bidding at their bankruptcy proceedings, she proposed, it might be possible to obtain detailed regulatory filings, manufacturing strategies, and safety data—types of information usually considered trade secrets. She called these documents “biotech’s lost archive” and said they could be used to help train AIs that would act as powerful copilots in the often opaque drug approval process. 

Today the OpenAI Foundation, the nonprofit parent of OpenAI, said it would fund her idea as part of a new effort it calls Public Data for Health, which aims to help artificial intelligence make big leaps in medicine by paying to create “high-quality scientific datasets.”

The basic idea is that AI isn’t going to be capable of making important breakthroughs in curing disease unless researchers can feed the models much more information than they have so far. 

“Everyone is recognizing that data is the biggest bottleneck in successfully applying AI to biology,” says Morgan Levine, a former vice president for computation at Altos Labs, a longevity company.

In its initial round of data grants, the OpenAI Foundation also announced that it would give $40 million to a program to collect data about novel cancer vaccines at the University of North Carolina, Chapel Hill, and support OpenAdmet, a group that runs competitions in which researchers try to predict drug effects. 

Teslo’s idea for a biotech archive received $500,000 and will be pursued by 1Day Sooner, an advocacy group representing clinical trial volunteers, which she advises.

“We expect many remaining breakthroughs in preventing and curing disease to come from pairing the intelligence of new models with more observations of the world—in other words, more data,” the OpenAI Foundation said in a statement.

OpenAI started as a nonprofit, but leader Sam Altman restructured it to form a for-profit corporation that develops new models, launches products, and is now planning an initial public offering of stock that could value it at $1 trillion.

Because the foundation holds a 26% equity stake in OpenAI, it is now be on track to become the richest charitable organization on the planet, potentially sitting on $250 billion in stock value. (By comparison, the Gates Foundation and a trust associated with it held about $180 billion at the end of 2025.)  

Making good use of that kind of money will not be easy. The foundation, based in San Francisco, is still hiring for many key roles and started ramping up its grantmaking only this year. Its largest single gift so far, of $100 million, was awarded in August to the Common Health Coalition, an organization that helps patients get access to drugs for hepatitis C.

OpenAI’s charitable efforts come even as apocalyptic fears have broken out about the possibility that runaway AI could wipe out all human life, possibly by launching a deadly bioweapon.

Those fears have been stoked by AI company insiders, some of whom say the chance of human extinction within the next decade is 10% or more. Last week, Altman and xAI founder Elon Musk both endorsed a call by Anthropic CEO Dario Amodei to “slow the pace at which we improve the capabilities of AI models” so that risk prevention can catch up.

Jacob Trefethen, an executive at the foundation, says it essentially operates separately from OpenAI but shares an official mission of ensuring that artificial intelligence “benefits all of humanity.”

“We’re starting grantmaking when we think the best way to achieve that mission is to make grants to external nonprofits, research institutions, and other third parties,” Trefethen said in an interview. He says the foundation hopes to give away $1 billion by the end of the year. 

The $500,000 grant to 1Day Sooner will help the group prove it can obtain the data troves of bankrupt companies, says the organization’s president and cofounder, Josh Morrison. He thinks nonexclusive copies of company datasets could be acquired for only “a few tens of thousands of dollars” each.

His organization is currently in possession of three datasets, two of them donated by Lumen Bioscience, a biotech that previously used the Chapter 11 strategy to gain insights into another company’s drug development efforts. 

Morrison says two other attempts to obtain drug company files this year proved unsuccessful, after 1Day Sooner’s bids were not accepted. 

Bankruptcies could become what some are calling a “new land grab” for AI training. Last month, Google won a bid to take over the corporate data of the failed carrier Spirit Airlines, including 100 million emails. That led to objections from flight attendants and others who worried that private or proprietary data could be exposed. 

The drug company files that 1Day Sooner is seeking are known as common technical documents. They typically contain the back-and-forth between companies and regulators, as well as detailed scientific and medical measurements, and essentially provide everything that is known about a drug.

According to Teslo, who is a writer for Works In Progress and a nonresident fellow at the Institute for Progress, a think tank in Washington, DC, a stockpile of such files could help turn an AI into a regulatory expert, which in her view could be one of the main ways AI helps speed cures to market.

“People say ‘We will invent AI, and AI will cure cancer,’ but that’s very removed from the messy reality and the regulatory process,” she says. “About 70% of the money and time in drug development is spent in clinical development—organizing the trials and testing the drug—but despite that, the process is basically a black box, especially for small biotech companies generating the innovations.” 

Read the full original article:

MIT Technology Review
#OpenAI