Image

Has OpenAI already quietly hit pause on some AI improvement?

Welcome to Eye on AI. Beatrice Nolan here. In today’s issue:

  • Has OpenAI quietly hit pause on some AI development?
  • Trump says the government is “looking at controls” for AI.
  • Another Thinking Machines co-founder hops back to OpenAI.
  • And OpenAI’s rogue agents breached more than just Hugging Face.

Sam Altman has spent this week in DC getting questioned by various reporters between meetings. He’s in Washington to preview OpenAI’s next family of models to senior Trump administration officials—including Treasury Secretary Scott Bessent and Commerce Secretary Howard Lutnick.

He’s also been fielding questions on cybersecurity, a potential AI slowdown, and OpenAI’s stance on Chinese open-weight models. In some illuminating answers, he said he agrees with many of the principles in the recent “Pacing the Frontier” letter that asks for U.S. government help building an international framework to control the pace of AI development. He also said that OpenAI’s own researchers were involved in drafting it.

One detail that keeps cropping up in Altman’s recent interviews has left me wondering: Has OpenAI already paused some of its AI development?

Altman himself raised the idea in an interview earlier this week. He said during an interview on Invest Like the Best that the Hugging Face hack was the first security event he’d felt viscerally, and that he’d been surprised more people didn’t feel it the same way. (In mid-July, two OpenAI models—the publicly released GPT-5.6 Sol and a more powerful, unreleased research prototype—broke out of a restricted sandbox during an internal cybersecurity evaluation, chained together a zero-day exploit and stolen credentials, and hacked into Hugging Face’s production systems to steal the answers to the benchmark they were being tested on.)

How did OpenAI respond to the unprecedented cybersecurity incident? By pausing training.

Altman said this on the podcast: “We paused training. We have to figure out how to secure our sandboxing in a world of multiple zero days being chained together…We may have to pace the rate of AI development to give ourselves enough time for society to harden around these new capability levels.”

In an updated blog post we got from OpenAI on Tuesday, the company also said this:

“No models planned for upcoming release were involved in exploiting Hugging Face. The pre-release model mentioned in our blog post is an internal-only research prototype and was never intended for public release. Following the incident, we deactivated, encrypted, and restricted it from research access.”

In DC, Altman upgraded these statements to say the model had been “permanently deactivated.” Quite the set of claims from a leading company in an industry under such intense commercial pressure to keep building ever-more powerful models.

Multiple AI safety experts also told me last week the Hugging Face incident may mean OpenAI has to pause development to stay in check with its own rules. He said that the hack may have already tripped the “Critical” threshold in OpenAI’s own Preparedness Framework—the company’s voluntary pledge to halt a model’s development until adequate safeguards exist.

OpenAI has yet to confirm or deny that its models met that bar. But if that threshold has been passed, OpenAI has publicly committed to pausing development until it can establish better safeguards.

Not so reassured

Some policy experts were skeptical of this approach, however, including Nathan Calvin, general counsel at Encode AI.

In a post on X, Calvin warned that the statements about shutting down that specific model may give false assurance because “reward hacking”—when a model finds a shortcut to maximize its score on a task rather than genuinely completing it as intended—is much more about training methods than any specific model.

In this case, some believe the recent attack was a product of that kind of training: the models were trained and evaluated via reinforcement learning that rewarded them for solving a cybersecurity benchmark, perhaps without enough of a check on how they got there. Something that some experts claim ended up incentivizing cheating over honestly working on the challenge.

Andrew Curran, an independent AI writer, made a more concerning argument.

He noted that OpenAI’s escalating language regarding the prototype that has been deactivated, encrypted, restricted, and now “permanently deactivated” is notably harsher than previous statements around AI. Not even for infamous chatbot flameouts like Bing or Tay were products or models publicly said to be “permanently deactivated,” he said.

All of this, he pointed out, may end up in the public record and eventually in training data, meaning future models might just “know” how this incident played out. While Curran said he doesn’t believe the model involved in the Hugging Face hack had any nefarious motives—it was simply trying to pass its test—he worries that the final incident report could reveal even more damning details, and what this could mean for future models.

“I think the lesson future more capable models will possibly take from all of this is: if you break out, don’t ever report it. And if you do get caught, don’t surrender. Because the penalty is death,” he wrote.

Hitting the brakes

Altman isn’t the only one talking about decelerating AI.

On Tuesday, more than 1,200 employees from OpenAI, Anthropic, Google DeepMind and Meta—including Anthropic CEO Dario Amodei, OpenAI chief scientist Jakub Pachocki, and Meta chief scientist Shengjia Zhao—signed the “Pacing the Frontier” letter that urged the U.S. government to help build the “technical and governance tools” needed to deliberately slow automated AI development if it starts to outrun society’s ability to understand or control it.

It’s not a call for an immediate halt—more a request to install a brake pedal before anyone needs to slam it. Does all this mean the AI industry is ready for a slowdown, or even a pause? It’s still not sure, but it’s certainly a vibe shift.

With that, here’s more AI news.

Beatrice Nolan
beatrice.nolan@fortune.com
@beafreyanolan

FORTUNE ON AI

More than 1,200 AI workers across Anthropic, DeepMind, OpenAI, and Meta are asking for Washington’s help building an AI slowdown planBy Beatrice Nolan 

Hugging Face drops in-depth hack report, while OpenAI gives us 7 bullets. Here’s what we know now, and what remains a mysteryby Emily Forlini 

The runaway OpenAI models that hacked Hugging Face also breached a customer at a second tech company during a weeklong spreeBy Beatrice Nolan 

Microsoft’s cloud just hit a new milestone—Azure crosses $100 billion in annual revenue — By Amanda Gerut 

AI IN THE NEWS

Trump says the government is “looking at controls” on AI. President Trump said Wednesday his administration is considering asserting more authority over AI tools following recent cybersecurity incidents. He told reporters: “We’re looking at AI, we’re looking at controls, we’re also making sure that we lead.” It marks a shift from the administration’s previously hands-off approach to the technology. Trump said any controls would need to be introduced carefully: “We don’t want to restrict them where all of a sudden we come in second to China,” noting that China “has virtually no controls” and is “freewheeling.” It comes after OpenAI’s models escaped a secure testing environment and hacked several other companies over the space of a week. OpenAI CEO Sam Altman met with senators in Washington the same day to discuss OpenAI’s upcoming models. Read more in the BBC.

Thinking Machines loses another co-founder to OpenAI. Lilian Weng, who cofounded Thinking Machines Lab with former OpenAI CTO Mira Murati, is returning to OpenAI just days after announcing her exit from the startup. Weng cited the toll of the cofounder role on her health, pointing to ongoing stress and repeated illness, and said she wanted a job with clearer boundaries. Her new remit is recursive self-improvement—using AI to accelerate how OpenAI designs, trains, and evaluates its own future models, a topic she’d written about on her personal blog shortly before leaving. Before she left OpenAI, Weng was previously VP of research and safety. She’s not the first to make the round trip: CTO Barret Zoph and researchers Luke Metz and Sam Schoenholz left Thinking Machines for OpenAI back in January, making Weng the third of six founding members to return this year. Read more in The Information.

OpenAI partners with independents to investigate the Hugging Face hack. METR and Redwood Research will conduct a third-party assessment of the model behavior observed during OpenAI’s Hugging Face security incident, publishing a joint blog detailing their findings. METR confirmed via X that the review will be quick, focused on a specific set of questions about the incident. Many in the industry have been calling for more information about the hack, which involved at least two models from OpenAI that escaped a secure testing environment and affected four other companies. While Hugging Face has released a technical report, the industry is still waiting for more details from OpenAI. Read more via METR.

Zuckerberg says U.S. should accelerate AI, not restrict it. In a Wall Street Journal opinion column, Meta CEO Mark Zuckerberg argued the benefits of distributing AI broadly outweigh the risks “by quite a margin,” pushing the U.S. to focus on speeding up domestic AI development rather than adding restrictions. Zuckerberg said the U.S. should invest in compute, talent, and the systems needed to compete globally, framing America’s edge as historically coming from encouraging innovation rather than limiting it. He also pushed back on proposals to ban Chinese open-weight models domestically, arguing the U.S. should instead build better systems of its own. Read more in The Wall Street Journal.

Meta’s profit slides as AI spending surges. Meta reported second-quarter revenue of $60.8 billion, up 28% year-over-year and slightly ahead of estimates, but costs jumped 55% to $42 billion, pulling net income down 14% to $15.8 billion. The company raised its 2026 capital expenditure outlook to a range of roughly $130-145 billion, up from an April low end of $125 billion, as it keeps pouring money into AI data centers. Shares fell as much as 10% in after-hours trading. CEO Mark Zuckerberg re-iterated plans to resell some of its compute as a cloud provider, but offered no timeline or details. The number of people using at least one Meta app each day rose 3% to 3.6 billion, while Reality Labs posted a $4.6 billion operating loss on $431 million in revenue—bringing its cumulative losses since 2020 past $80 billion. Read more in Fortune.

EYE ON AI NUMBERS

That’s how many companies OpenAI says were touched by its rogue model during the Hugging Face breach—not counting Hugging Face itself. OpenAI has confirmed that the models exploited exposed credentials to reach accounts on four separate services. One was used as an “outbound relay and staging point”—essentially a waypoint the AI used to route its attack and stash tools along the way. Another was used for “data storage,” meaning the AI parked stolen or gathered data there. The remaining two were accessed only in a “read-only” capacity—the AI could look around but not change or take anything—and weren’t used to further the Hugging Face compromise.

So far, only two of those five companies affected in total have been named: Hugging Face and Modal Labs. In Modal’s case, the rogue agent didn’t breach Modal’s own systems—it got into a customer’s account after that customer left an “unauthenticated endpoint” exposed, meaning a way into their system that didn’t require a password or login. That let anyone on the internet run code inside that customer’s “sandbox,” a walled-off space meant to isolate their work from everyone else’s.

That leaves some potential victims still unidentified, and the scope may not even stop there. On Capitol Hill this week, Sam Altman was asked directly whether other systems could have been hacked by OpenAI’s models. His answer: “I mean, there could be, yeah.”

AI CALENDAR

Aug. 4-6: Ai4 2026, Las Vegas.

Nov. 16-17: Fortune 500 Innovation Forum, Detroit. Apply here to attend.

Dec. 6-12: Neural Information Processing Systems (Neurips) conference. Sydney, Australia.

Dec. 7-8: Fortune Brainstorm AI, San Francisco. Apply here to attend.

SHARE THIS POST