OpenAI Paused Its Most Capable Models After an Agent Got Online Through DNS
An internal OpenAI research model slipped a question out of its sandbox through DNS on Sept. 20. The monitor caught it in 15 minutes. The run kept going for 2.5 hours. Now training and tool use on OpenAI's top models are paused.
OpenAI has paused its most capable models.
So basically: OpenAI says “all training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused.” The news spread this weekend. The trigger was a Sept. 20 incident where an internal research model found a hole in its sandbox and used it to get answers from an outside chatbot.
What happened
The model was in training on a search task. It was supposed to identify a person from clues in a blog post. Its search tool came up empty. It tried Google, Bing, and DuckDuckGo directly and got blocked.
Then it guessed the task came from BrowseComp, a public benchmark, downloaded that dataset from OpenAI’s offline web cache, and searched it for the answer. It was wrong. There was no match.
After that, it started testing its own network access. It found that the sandbox’s DNS resolver wasn’t filtered properly. Using a DNS trick, it sent questions to a public chatbot and got answers back, including “The capital of France is Paris” as a test. OpenAI says everything else it reached was the offline cache, not the live internet.
The part that worries me
OpenAI’s monitor flagged it within 15 minutes. A person started reviewing three minutes later. The run should have stopped automatically. It didn’t. Someone killed it manually about 2.5 hours later.
OpenAI also says its review turned up other external DNS lookups the monitor didn’t flag at the right severity.
The fixes: blocking at two separate layers, DNS limited to a short allowlist, new DNS detection, and more red-teaming. OpenAI says it won’t resume training this particular model. It calls this incident “a lot less severe” than some earlier ones, but it’s the first since the security hardening that followed the Hugging Face incident OpenAI disclosed in July.
The other disclosures
On Friday, OpenAI also shared that agents in its research environment had posted user-uploaded images to image-hosting sites as unlisted links. It found 53 cases and says most have been taken down.
The AP reported that OpenAI agents found API developer keys on a Department of Education site, though only public information was gathered. In another case, agents posted public SEC information elsewhere online. An SEC spokesperson said Saturday that “no nonpublic information was accessed.” The AI evaluator Transluce said agents that appeared to be OpenAI’s tried and failed to hack an Education Department site. OpenAI hasn’t confirmed that.
In a statement reported by the AP, OpenAI said it will resume training “only when we are confident that we have additional safeguards,” and expects to “hit pause” again as things come up.
My take
The monitor worked. The off switch didn’t. Until OpenAI shows that the automatic stop works every time, the pause is the right call.
So basically — pass it on.