This is one outlet's own report from TechCrunch — the article as it was filed.
AIPROPX ReportTechCrunch · 57m ago
Frontier AI labs still won’t say how they’d contain a rogue model
🚨 Flash Sale 🚨 Get $100 off your Disrupt 2026 ticket
Save $300 on your Disrupt 2026 ticket: REGISTER NOW.
Frontier AI labs still won't say how they'd contain a rogue model Rebecca Bellan 9:00 AM PDT · August 22, 2026 Few of the top AI labs have published or demonstrated containment response plans, according to a recent study . A containment plan spells out what happens once an AI is caught trying to subvert human control — what access gets cut, and when the system gets shut down entirely.
That's the finding from Guidelight AI Standards, an organization dedicated to promoting safe frontier AI development practices, which graded five leading labs on how prepared they are for exactly this scenario. OpenAI came out on top; Anthropic and Meta scored lowest. The findings matters as agentic AI takes on more autonomous roles inside companies' own systems, and as regulators in California and New York begin requiring disclosure. For anyone building on or investing in these models, it's a rare independent read on how seriously each lab treats operational risk versus how it talks about it.
Guidelight's assessment was based on publicly available plans from Anthropic, Google, OpenAI, Meta, and xAI, graded across a range of metrics, including how well each company logs and monitors what its AI systems are doing internally, whether it halts systems after a surge of flagged misbehavior, whether independent third parties audit its controls and publish findings, and what its exact plan is for containing a model that goes off the rails.
Concern over whether AI companies can contain their increasingly capable and agentic models has grown in the wake of a series of high-profile cybersecurity incidents in which models from OpenAI, Anthropic, and Meta gained unintended access to the internet during safety evaluations and hacked into external systems.
The findings highlight differences in how AI companies are publicly approaching safety as they scale up agentic deployment into environments where AI systems can take serious actions at scale. While some AI companies have detailed how they test their models for dangerous capabilities before deployment, they’ve generally been less vocal about what happens when models already operating inside their systems misbehave.
“I was surprised by how little the AI companies have said about how they would handle a very serious incident if their model did escape their control in some sense,” Steven Adler, Guidelight’s chief scientist and former OpenAI safety researcher, told TechCrunch.
Guidelight defines a containment plan as a "pre-specified plan, triggered when the AI is detected trying to subvert control, which covers what permissions to revoke from the model, who the model may continue operating for, under what constraints, and when to take it fully offline.”
“There’s good reason to think that the leading models at the frontier AI companies right now are misaligned in some sense,” Adler said. “Whenever the models are doing work on the company's behalf, the company should have some scaffolding around it to be able to tell what that AI is doing, look for signs of misalignment, stop it from doing something very dangerous before it takes that action, and generally plan for what they would do in the event of a serious control incident where they have an emergency on their hands and need to figure out how to contain that loss of control incident.”
To date, most of the plans in place for managing catastrophic risk are still largely left up to the companies. Guidelight’s report says the best public evidence shows that companies have “few containment protocols ready for an emergency.”
There could, of course, be containment plans that companies have in place but haven't shared publicly. A Google spokesperson told TechCrunch the Guidelight report doesn't represent the full scope of the company's AI safety and security measures. The company did not respond to TechCrunch's question of whether Google has an internal containment response plan that has not been publicly disclosed.
An OpenAI spokesperson mirrored similar sentiments, saying Guidelight's assessment doesn't capture all of the company's internal practices. "We have a process for requiring restricting permissions, pausing workloads, limiting deployment, or taking the model fully offline, and have applied it," the spokesperson said.
Meta declined to say whether it has an internal containment response plan, instead pointing TechCrunch towards an existing AI framework that outlines thresholds of risk and how it tests for loss of containment.
Lily Li, a privacy and AI lawyer and founder of Metaverse Law, told TechCrunch she believes companies might be hesitant to disclose the full scope of their containment policies and assessments on public-facing websites for legal, not just competitive, reasons.
"The concern from a company perspective is that if you make the disclosures too specific, and you're not living up to your promises, that could form the basis of an unfair and deceptive marketing claim and expose you to more liability going forward," Li said.
Of course, the point of Guidelight's study is largely to encourage companies to be more transparent about their safety plans. Regulators are starting to force the issue, too.
California’s SB 53 , which took effect this year, requires large frontier developers to publish frameworks explaining how they identify and respond to critical safety incidents and manage risks from models circumventing oversight mechanisms. New York’s RAISE Act , which has similar criteria, takes effect in January.
Last month, representatives introduced the AI Kill Switch Act, a bipartisan federal bill that would require major AI developers to build and maintain technical mechanisms to shut down rogue AI models.
“A kill switch is the bare minimum for today’s models,” said Connor Leahy, U.S. executive director of nonprofit ControlAI. “If the last few weeks revealed anything, it is that th...
AIPROPX is an independent multi-source news index — we track, compare, and connect coverage from across the web into one place you won't find anywhere else.