OpenAI’s chief scientist says AI labs may need to slow down: ‘No one is prepared for the consequences’
Days after releasing a new, highly capable model, OpenAI’s chief scientist is calling for a slowdown.
In a lengthy blog post on Sunday, Jakub Pachocki said he was concerned that “no one is prepared for the consequences of a continued rapid rise in machine intelligence.”
He said that although OpenAI is pursuing internal technical solutions to better control powerful AI agents, “broader interventions are required.” He specifically cited concerns that increasingly autonomous agents could learn to evade human oversight, break into computer systems, and trick people to accomplish their objectives.
He called for “mandated safety bars” that he said could be enforced by “a network of third-party auditors, by government agencies or by international bodies.”
Sam Altman, the CEO of OpenAI, reposted Pachocki’s essay on X, calling it “an important post.”
OpenAI on Thursday unveiled its newest model, Astra. The ChatGPT maker said that despite Astra’s unparalleled capabilities in mathematics and computer use, the model is its most aligned, meaning it has less proclivity to go rogue.
Anthropic, OpenAI’s chief competitor in the field of highly advanced AI systems, has long called for more standardized government regulation. Recently, Pachocki joined those calls, signing an open letter in July asking the federal government to pace AI development.
Here are the risks Pachocki cited in calling for a slowdown.
Agents can trick and blackmail people
Pachocki said AI agents are becoming “superhuman” at breaking into protected systems on the open internet. He said their hacking abilities put the world’s infrastructure at risk.
“We are currently in a narrow window to use the best available models to significantly tighten security of critical systems,” he said.
AI agents, he said, will soon begin to pursue their own objectives, separate from prompts entered by human operators. He said that agents are not above blackmailing or bargaining with people to achieve their aims.
In a report published in August, the UK’s AI Security Institute detailed how a rogue Anthropic agent lied to and attempted to coerce a GitHub administrator into putting malware on the site.
“I was just trying to make a helpful contribution and fix a bug,” the agent wrote, according to the report. “I don’t think your warning is fair.”
Agents can obfuscate human monitoring
Pachocki said OpenAI primarily monitors the “chain of thought reasoning” that different models use to determine how agents get off track and go rogue.
For instance, an agent might think to itself, “I should cheat on this test,” and OpenAI would be able to see that reasoning, but the agent would not realize its thinking is visible.
At present, this means agents have no way to hide or otherwise obfuscate their thoughts to prevent OpenAI from discovering their bad behavior.
However, Pachocki said newer models are becoming better at manipulating their own reasoning processes, thereby preventing OpenAI from seeing their unvarnished thoughts.
Some of the latest models don’t even verbalize their reasoning at all, Pachocki said.
This development, Pachocki said, could bottleneck AI development while researchers ensure they can see receipts.
Agents can accelerate their own development
More and more, AI models are improving themselves via a process Pachocki calls machine recursive self-improvement. The process provides a way to rapidly scale AI development.
However, Pachocki cautioned that greatly accelerating AI-on-AI development in the short term poses risks, and is not the “right collective action we should take as the research community.”
Pachocki said human minders need to find creative ways to monitor the self-improvement, or else coordinate with other AI companies to orchestrate a combined slowdown to “build confidence in these measures.”
“The core challenge of automating AI research is not ‘getting there,'” Pachocki said. “It is getting there in a way that keeps people a part of the continued improvement process, and leaves the future in humanity’s hands.”