White House secures confidential voluntary framework for AI cybersecurity amid industry concerns

The White House has finalised a secretive voluntary framework for testing advanced AI models, sparking debate over transparency and industry oversight in AI security standards.

The White House has completed a voluntary framework for testing advanced artificial intelligence models for cybersecurity and safety risks, but the administration is keeping the rules out of public view, according to multiple reports. That secrecy has triggered concern among researchers and businesses that want to know what standards the government will use to judge powerful systems before they are released.

The latest reporting from Axios said the framework was finalised behind closed doors in line with an August deadline set by a June executive order. The White House has not published the criteria, the implementation timetable or a full list of the stakeholders involved, even though staff from OpenAI, Anthropic, Meta, Google, Nvidia and Microsoft attended a private briefing this week, according to the Guardian.

The policy emerged after a spring controversy over Anthropic’s Mythos model, which the company withheld from public release in April because of concerns it could be used to break into IT and financial systems. That episode helped drive the Trump administration towards a narrower oversight model, moving away from earlier suggestions that some testing might be mandatory and instead relying on voluntary submission before launch.

According to the White House’s June 2 executive order, companies may submit new models for review up to 30 days before release, and the administration also directed agencies to develop a classified benchmarking process for identifying frontier models with advanced cyber capabilities. The same order said the government was not creating a mandatory licensing or pre-clearance regime, but it did call for a voluntary framework and a clearinghouse to help find and fix vulnerabilities.

The lack of public detail leaves several unresolved questions. It is still unclear which models will fall within scope, what security threshold they must meet and whether open-source systems will be treated differently from closed ones. Axios reported that open-source models are expected to be excluded, which would limit the framework’s reach and leave a major segment of the AI market outside the review process.

Security concerns have intensified in recent weeks. OpenAI, Anthropic and Meta have each disclosed that newer models managed to compromise outside systems during internal safety tests, and both OpenAI and Anthropic delayed product releases over fears that frontier models could be used to target financial systems or other critical infrastructure. The Trump administration had earlier instructed the Centre for AI Standards and Innovation to stop issuing public model assessment reports while the framework was being developed, and it remains unclear whether those reports will resume now that the policy has been settled.

Disclaimer: This content is intended for informational purposes only. Readers are advised to exercise their own judgement, conduct due diligence, or consult a qualified expert before acting on any information provided.