This is one outlet's own report from The Register — the article as it was filed. Other outlets are covering the same event; open the full story to compare every source side by side.
KETTLE So, an OpenAI model broke out of its sandbox last week, made its way to the internet, then hacked its way into Hugging Face, stealing some internal data and credentials in the process.
You can listen to the latest episode of The Kettle right here on this page, as well as on Spotify , Apple Music , or YouTube where you can subscribe to get notified of the latest episode.
That's big news in the world of AI, but as El Reg cybersecurity editor Jessica Lyons and senior reporter Tom Claburn tell Kettle host Brandon Vigliarolo , it's not really the end of the world as we know it. Sure, it means there's some capable models out there, and maybe there's more risk from them than some might think, but the OpenAI/Hugging Face mess only happened because of some very specific circumstances. That, and it's actually a really good reason to prioritize more open models instead of relying on frontier labs to own the entire space.
Brandon: Hey everyone, welcome to another episode of The Register's Kettle podcast. Though honestly, maybe we ought to start just calling it The Reg Talks AI because, yet again, we're focusing on artificial intelligence. If you've been following the news in that space this week, you probably know what we're gonna be covering as there's no hotter topic in AI land right now than the fact that some autonomous OpenAI agents broke out of their sandbox and attacked AI model host Hugging Face, as the company admitted on Tuesday.
With me to discuss this breakthrough in AI threat capability is our cybersecurity editor, Jessica Lyons, and senior reporter Tom Claburn. Both have been on top of this. So thanks for joining me, guys.
Brandon: Yeah. So let's jump right into it. Jess, what exactly happened here? Let's start from last week when Hugging Face said it was attacked.
Jessica: Right, so Hugging Face disclosed that there had been a digital intrusion, and they said it was "driven end-to-end by an autonomous AI agent system." So these agents attacked a limited set of their internal datasets and then also credentials used by their services. So when they disclosed this, they didn't say or they didn't know which models had powered the agents.
They did say, though, that they tried to use these commercial models for the investigation, but the guardrails put in place, the safety guardrails, blocked the frontier models from actually helping them with the investigation. And because of that, they turned to a Chinese open-weight model, and that's how they discovered this agent swarm that had attacked some of their datasets and their production.
Brandon: OK, they didn't mention which frontier models they tested, did they?
Jessica: No. At the time they didn't. They said "we tried to use the commercial frontier models and they all refused because of their guardrails."
Brandon: Right. So probably trying to ask OpenAI models, hey, do you know who did this? We can't tell ya.
Jessica: Right. Exactly. That was kind of right. That was kind of the takeaway from all this. OpenAI is a Hugging Face partner. And so then that brings us to this earlier this week when OpenAI admitted that it was the operator of these agents that attacked Hugging Face. It said it was GPT 5.6 Sol and then "an even more capable pre-release model." Those were among the ones that attacked Hugging Face.
But it also said, and this was really important, that the models had their guardrails intentionally disabled because the whole point of this was to test for cyber vulnerabilities. So that's a big piece that seems to be missing in my opinion in a lot of the discussion here. And after OpenAI said that its models were involved in this autonomous attack, that's kind of when all hell broke loose and everybody said "this is what we've been warning about. There's autonomous agents attacking and they're not supposed to and the sky is falling."
Brandon: So, to be clear as to what happened with OpenAI, right? They were basically running some capture-the-flag exercises in a sandbox environment, right?
Brandon: Or something to that effect with their models and they disabled the guardrails so these things could basically use their full capabilities to try to solve these puzzles, right? And I think it was that they exploited a couple of zero-days to escape the sandbox?
And then they went after Hugging Face because they thought for some reason that Hugging Face may have solutions for these puzzles. Is that right?
Jessica: Right. So their prompt was to pursue advanced exploitation using complex attack paths. So that's what they were instructed to do, and that's exactly what they did. And it sounds like the models inferred that Hugging Face might have some ideas to help them actually do this. So the models essentially did what they were instructed to do.
Tom: One of the one of the things that didn't come up in their post is that OpenAI didn't seem to take any responsibility for "yeah, we should have been supervising this." That's, to me, the thing that really gets me is imagine Waymo saying "yeah, we conducted a test of our cars and we decided not to have any operators monitoring them remotely. We just let them go and we took away all of our safety guardrails and we're so sorry that it hit the kindergarten."
It's totally predictable that if you're gonna automate something and then not pay attention to it, you're gonna get unexpected results.
Brandon: Yeah, especially, like you said, with the safety guardrails all removed. You're literally asking for this potential thing to happen. I mean, obviously they probably didn't know there was some zero-day buried in something in the sandbox.
Jessica: It was exposed credentials and zero-days in the production database. And so that's how they got in. So it's not a crazy attack chain. The fact that agents found it is more notable, but it's not this super complex attack method.
AIPROPX is an independent multi-source news index — we track, compare, and connect coverage from across the web into one place you won't find anywhere else.