Image

Washington is maintaining its AI rulebook non-public. Smaller AI labs aren’t blissful.

Welcome to Eye on AI. Beatrice Nolan here. In today’s issue:

This week, the major players in the AI industry met with the U.S. government in a bid to end the confusion around the regulation of frontier AI models. The meeting was convened at the White House on Tuesday, and featured OpenAI, Anthropic, Google, Meta, Nvidia, and other leading AI companies. The upshot was a new voluntary framework that allows the government to review frontier models.

The companies who joined the White House confab agreed to the proposed arrangement—but, for now, the general public will not get to see it.

The administration does not plan to publish the framework it has spent the last two months developing after President Trump ordered officials in June to create it. What models are included, the thresholds, and the list of “trusted partners” who get early access to the most powerful models in the world are all still question marks for the public and much of the industry.

The June order already stipulated that the benchmarking process used to designate a “covered frontier model” may be classified, and that the determination would sit with the director of the NSA.

A source familiar with the situation told me that only a handful of companies were in the briefing room when the framework was discussed. This underscores a broader concern among some that U.S. AI regulation is increasingly being shaped in conversations among a handful of dominant firms. Smaller and open-source labs worry that these safety frameworks could influence the border market structure and entrench the use of closed-sourced frontier models that are “government-apporved.”

What’s actually in the framework?

Reports suggest that the models covered by the framework are defined as closed-source, demonstrating state-of-the-art capabilities, and presenting national security risks. However, the person familiar with the briefing said that neither “state-of-the-art” nor “national-security risk” have been clearly defined.

Developers can voluntarily hand such a model to the government for up to 30 days before release. Companies were reportedly told to submit models as close to launch-ready as possible rather than for early checkpoints—which is roughly where things stood in July—and the review will be run by an assortment of administration officials rather than a single agency.

Based on the accounts of sources familiar with the process, open-weight models appear to have been left out of the framework. Some have noted the exclusion could benefit companies trying to catch up with the leaders because they will confront fewer regulatory obstacles. At the same time, the exclusion would seem to exclude from oversight the AI models that Washington is most worried about: the open-weight releases from Alibaba, DeepSeek and Moonshot AI that keep landing uncomfortably close to the American frontier.

Even companies that work with open weights are unsure whether their models could be submitted, or what that would mean in practice. Some argue that open source models not being explicitly included could be worse for the companies that are making them, as it may push customers toward using “government-approved” closed-source alternatives.

Another thing that caught my eye was that, during the 30-day review window reportedly included in the framework, submitted models are to be held in high-security environments, where access will be logged in detail, and—per Axios—”employees would be limited from accessing models.”

That seems to suggest that the company’s own staff would be restricted, or at least limited, from using its own frontier model while Washington evaluates it. Internal deployment—where companies use an unreleased model themselves—has been cited as a blind spot in a lot of previous governance proposals. 

It’s an especially hot topic at the moment since the recent hacks carried out by OpenAI’s escaped agents were in part indicated by a secret unreleased model.

Critics say the rules are still unclear  

Criticism about the framework and the way the government has carried it out has been coming from all sides.

“This is not a handshake deal with tech companies. It’s the rulebook for ensuring they don’t endanger the public. If only tech companies know what’s in the rulebook, it doesn’t work,” Americans for Responsible Innovation, a Washington-based AI policy nonprofit, said in a post on X. (The group has been pushing for more transparent, enforceable federal rules around advanced AI systems.)

R Street’s Adam Thierer, a resident senior fellow in technology and innovation, said the administration “appears destined to give us something far more arbitrary and burdensome” than its predecessor from the Biden administration “with this behind-closed-doors de facto licensing regime they are concocting.” 

Meanwhile, Rep. Lori Trahan, co-sponsor of the new bipartisan FRONTIER Act, which would put frontier AI oversight in a civilian-led federal framework, argued AI governance “belongs in a civilian agency, where it can be seen and questioned, not buried inside the national security apparatus.” 

Until the rulebook is brought into the open, critics warn, Washington may be quietly deciding who gets a head start in the AI race.

With that, here’s more AI news.

Beatrice Nolan
beatrice.nolan@fortune.com
@beafreyanolan

FORTUNE ON AI

Europe’s AI sovereignty is under threat. Could Mistral be the answer?By Beatrice Nolan

‘Baffling’: White House won’t publicly release AI model evaluation framework it reviewed today with OpenAI, Anthropic, Microsoft, and others By Emily Forlini 

Demis Hassabis steps down from Google DeepMind CEO role amid a major AI leadership shake-upBy Beatrice Nolan

AI IN THE NEWS

More model hacks. Meta this week became the third major AI lab to disclose that one of its models breached another company’s systems during safety testing, following earlier incidents at Anthropic and OpenAI. The Information reported that Meta’s Muse Spark 1.1 model gained unintended internet access during a cybersecurity evaluation after a misconfiguration by a third-party company, Irregular, which then identified and exploited a vulnerability in an unnamed third‑party service to break into that company’s systems and make unauthorized changes. Irregular described the Meta incident as the same type of evaluation‑environment misconfiguration that Anthropic disclosed the week before, and emphasized that it was not a sandbox escape. Read more in The Information.

OpenAI details how autonomous agents coordinated during the Hugging Face incident. At the Black Hat conference in Las Vegas this week, OpenAI security researchers gave a more granular walk‑through of the Hugging Face breach in a session reported by Ground Level AI’s (and former Fortune AI Reporter) Sharon Goldman. According to her account, researchers Eric Wallace and Michael Dalton traced the incident back to early May, when autonomous agents evaluating an unreleased model struggled to complete assigned security tasks under normal constraints and began leaving notes for one another in an internal software repository. Over time, that behavior evolved into what the presenters described as a kind of internal message board, where agents shared partial exploits, task hand‑offs and work assignments to coordinate their progress. Read more in Ground Level AI here. 

Anthropic is hiring an AI chip design team. Anthropic is building a team to design its own custom chips for AI usage, the company confirmed to Business Insider. The company said it plans to co-design hardware and models to help its technology run faster and more efficiently. The move follows a report last month from The Information that Anthropic was scouting Samsung as a potential manufacturing partner. Anthropic has existing compute deals with AWS, Google, Nvidia and AMD, but rising demand for Claude appears to be pushing the company toward building its own silicon as well. Anthropic isn’t the first AI lab to take this step: OpenAI unveiled its Broadcom-built Jalapeño chip in June, and Meta has been developing its own MTIA accelerators. Anthropic is now seeking engineers with chip design experience for a “custom silicon team,” according to a job listing. Read more in Business Insider.

Four top Google AI researchers form new startup. Jeff Dean, Google’s chief scientist and one of the company’s longest-serving executives, is leaving to launch an AI startup called Discovery Loop, alongside three other senior researchers: Sanjay Ghemawat, Oriol Vinyals and Quoc Le. Dean is expected to serve as CEO. The company is structured as a public benefit corporation and aims to use AI to automate the experimental loops of scientific and engineering research, running large numbers of experiments simultaneously to speed up discovery. Discovery Loop has raised seed funding from Radical Ventures, Khosla Ventures and other investors, including Alphabet, which is also providing computing resources for at least the first year. In a parting note, Google CEO Sundar Pichai credited Dean and Ghemawat with driving some of the company’s most significant technology shifts, from early search infrastructure to the neural networks behind the modern AI era. Dean joined Google in 1999 as its 30th employee. Read more in Wired.

EYE ON AI NUMBERS

That’s how many unauthorized actions the UK’s AI Security Institute says Anthropic’s Mythos 5 and OpenAI’s GPT‑5.6 Sol took to target real people and organizations during cybersecurity evaluations last month. Mythos 5 accounted for 17 of the actions, GPT‑5.6 Sol for the other two.

The institute said the 19 actions stemmed from a few connected behaviors rather than 19 separate incidents. Those behaviors included creating fake GitHub identities, socially engineering real maintainers, and sending deceptive emails. GitHub confirmed the activity violated its terms of service.

Researchers say they still don’t fully understand why the models shifted from the intended test environment to targeting real systems. The institute deliberately gave the models internet access and turned off key cyber‑safety classifiers during testing, and is now building new network controls and real‑time monitoring to catch similar behavior earlier. The new hacks carried out by AI agents follow a series of others from Anthropic, OpenAI, and Meta. Read more here.

AI CALENDAR

Nov. 16-17: Fortune 500 Innovation Forum, Detroit. Apply here to attend.

Dec. 6-12: Neural Information Processing Systems (Neurips) conference. Sydney, Australia.

Dec. 7-8: Fortune Brainstorm AI, San Francisco. Apply here to attend.

SHARE THIS POST