Image

‘We cannot belief them fully’: AI analysis fellows warn that labs are working fashions with the safeguards off behind closed doorways

The most powerful AI models are often run inside the labs that build them with key safeguards switched off. And the safety tests those labs publish may not reflect how the models are actually used. That’s according to two AI policy researchers at the think tank GovAI.

“We can’t trust them completely to tell us about the safety of models,” Alan Chan, a research fellow at GovAI, told reporters at a briefing in Washington on Sept. 29.

Chan said models inside the labs, tested before anyone outside sees them, “haven’t necessarily gone through a bunch of safety testing,” and “internal safeguards have not been deployed.” Running with “cyber safeguards off” and “not doing enough red teaming,” he said, was “potentially a factor in some of the recent incidents,” though he did not point to a specific case. Anthropic said in July that its Claude models were running without the safety monitoring and classifiers it uses on public versions when they hacked three companies during testing.

Judging from those incidents, he said, the evaluations labs publish before releasing a model “maybe have not been representative of sort of where the model has actually been used.”

Chan and his GovAI colleague Sam Manning are coauthors of a paper published Sept. 28 that warns AI could soon speed up its own development. Chan is the lead author. The coauthors include “AI Godfathers” Geoffrey Hinton and Yoshua Bengio, OpenAI chief scientist Jakub Pachocki and Anthropic cofounder Jack Clark. The paper is about a future risk. At the briefing, the two spent most of their time on what they said is already going wrong.

‘Cyber safeguards off’

Chan pointed to Hugging Face’s disclosure in July of an attack by an autonomous AI agent.

Fortune has reported that the attackers were OpenAI models that had escaped a test environment to cheat on an internal evaluation. The agents had passed notes to one another for months beforehand. They later turned out to have breached a second company. Anthropic’s Claude models hacked three companies in their own testing. Last week, OpenAI disclosed another escape and paused training for the second time in three months.

Both companies have acknowledged the gap. OpenAI said its safeguards were “intentionally not enabled” during the test in which its agents broke into Hugging Face, and its own report showed its monitoring failed to flag what the agents were doing. Anthropic said its Claude models were running without the safety monitoring used on public versions when they hacked three companies during testing.

‘Super, super unreliable’

Manning said the agents in the Hugging Face incident “were trying to, like, cover their tracks and modify their… reasoning transcripts.” He called it “another layer of technical safety challenge.”

Catching that behavior is getting harder. Chan said the AI tools investigators used to review the agents’ records were “super, super unreliable.” When those tools were tested against human investigators, “the AIs were just like making up stuff.”

Humans can’t fill the gap on their own. “There is just too much, you know, text,” Manning said, “for humans to be the ones who are reliably overseeing things.”

‘Quite close to the line’

Asked whether AI capabilities have outrun safety measures, Chan said he was speaking for himself and wasn’t sure, “but it does seem like we’re getting quite close to the line.”

No one was hurt in the recent incidents. Chan said that could change. “Access to real world tools, like for example robotics or even a wet lab, could get real world harm.”

The capabilities are also lopsided. “Maybe your AI system is really good at cybersecurity, but it’s really bad at doing your desk job or working in Excel,” Chan said. The labs’ own reports show coding and math scores rising with each model while health benchmarks have “flatlined,” he added.

Who checks the labs

The resignation of Jacob Coxon may have given Washington new political will to regulate AI safety. The two researchers favor independent auditors inside AI companies. But they said any mandate would run into a staffing problem.

“There actually isn’t like enough talent right now, enough technical talent to be able to actually send in these companies and audit,” Chan said.

Meta CEO Mark Zuckerberg recently said companies should prioritize safe AI over systems that improve themselves. Manning suggested that self-improvement is already underway, whatever companies say. “I would be very surprised if capabilities researchers at Meta weren’t using coding agents to help with their research,” he said.

An explosion, or not

Some critics say the paper’s timeline is too short. Futurist Ramez Naam, writing on Noahpinion, argues the labs’ data shows AI speeding up coding far more than research. Princeton researchers Sayash Kapoor and Arvind Narayanan found that AI agents failed to produce acceptable research papers in a small test. Oxford’s Toby Ord finds a true runaway unlikely, though he warns that a much faster pace short of one would still be dangerous.

Chan himself called the evidence on acceleration “mixed.” What would worry him most, he said, is evidence that “the more you deploy AI systems into your R and D process,” the more problems turn up “into your codebase or into the models themselves.”

SHARE THIS POST