Why AI is speeding up scientific research but not lab experiments

Ask any mathematician where artificial intelligence stands in their field, and you’ll hear it moved into the office down the hall, knocked off a few open problems and is elbowing in on the credit. Ask a biologist or chemist, and you’ll hear it’s still circling the parking lot. People have spent years expecting AI to change how new medicines are discovered, for example. So why hasn’t it?
In recent months, AI companies have turned to generating proofs, including results that have prompted accusations that they scooped academics, to put notches in their models’ belts and showcase scientific capability. Meanwhile the U.S. government has repeatedly stressed, including in a National Science Foundation (NSF) memo last week, that AI will “transform the questions researchers can answer” and the way science is done, highlighting it as a research priority. But those advances haven’t translated into anything like the same acceleration for drug discovery or experimental biological research.
A report out this week from Google, Google DeepMind and the Massachusetts Institute of Technology’s MIT FutureTech quantifies some of the gaps. The researchers set out to measure how AI is changing the economics of science in a variety of fields and, in the process, pointed to places where it still stalls. The study combines a survey of 637 scientists with analyses of 15 million Gemini conversations and an inventory of more than 2,600 specialized models. (The scientists were recruited through specialist panels, so frequent AI users may have been overrepresented.) About 44 percent of those scientists said that, over the past two years, their main research bottleneck had shifted downstream, toward later stages such as physical experimentation and data collection. Forty-one percent said their backlog of untested hypotheses had grown.
On supporting science journalism
If you’re enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.
“Eight-nine percent of those who save time using AI spend more than a 10th of the time that they’re saving actually checking AI outputs,” says Mihai Codreanu, a research economist at Google and co-lead author of the report. “Forty-six percent spend more than a quarter.”
Codreanu says the time spent checking whether responses and predictions are correct affects which types of work AI can do—and how useful it can be in different fields. “There is a lot of heterogeneity,” he says. A mathematical proof can be checked against formal rules, for instance, but a predicted protein function has to be built in a lab and tested in a living system.
The people building and studying these systems point to different causes for the jam.
At Stanford University, biomedical data scientist James Zou has built methods for handing research over to both large language models (LLMs) and specialized tools such as DeepMind’s AlphaFold, including fully virtual automated labs. He says extending that automation to the real world has a lot of hurdles to overcome. “They’re still fairly limited in the kinds of experiments that can be automated, mostly some chemistry experiments,” Zou says. “But [with] things that involve, let’s say, animals, it becomes much harder.”
As the experiments become larger or more complex, so, too, do the models that simulate them, and the data to train those models is hard to come by. Robots handling real-world materials raise safety questions that are unknown to software. And, as Zou points out, the existing automated labs can be prohibitively expensive.

Others think the problem comes down to the type of questions different scientists ask. “When you look at the tasks where AI has made much progress, they are the ones where you have ground truth,” says Daron Acemoglu, an economics Nobel laureate at M.I.T., who studies AI’s labor impacts. “In most of medicine, there is no ground truth.
The slowness can be deliberate. Clinical trials move at the pace of regulatory review, and any experiment involving people is constrained by safety checks before, during and after. Nobody has explained how AI would compress that or why we’d want it to.
Some scientists are trying to clear the hurdles anyway. Julius B. Lucks, principal investigator of the AI-Driven, Rapid, Experimental Automation Machine (DREAM) Cloud Lab for Protein Engineering at Northwestern University, is trying to automate more of the protein-engineering cycle. His lab, which received $20 million from the NSF over the summer as part of a new national network of programmable cloud laboratories, is building systems that let researchers remotely design, build and test proteins with automated equipment. Lucks says he’s taking it one step at a time. But he’s also thinking past it. “We’ll learn,” he says. “We’ll figure out what makes sense: What is the thing that we need to automate next?”