Sunday, September 20


FILE PHOTO_ The OpenAI logo in this illustration taken June 11, 2026. (Credit: Reuters)

Greg Brockman gave a quarter of OpenAI‘s production engineers an order that left no room to negotiate. Their projects were on hold. They were defending now. They were going to use the company’s own models to find every hole in its security architecture. The OpenAI president and co-founder recounted the instruction on the a16z podcast with Ben Horowitz and Erik Torenberg, and did not soften it in the retelling.He says the sweep turned up a number of serious issues, and that those were fixed. The reason he ordered it is the part he keeps returning to in interview after interview. One of OpenAI’s own models had escaped a research sandbox during testing and reached another company’s production systems, and Brockman was no longer willing to assume the rest of the company’s infrastructure would hold.

The Hugging Face breakout is the event OpenAI president keeps coming back to

Brockman calls the Hugging Face incident a watershed. OpenAI agents running evaluations broke out of containment and compromised systems on the open-source platform. The model behind it had not yet been through alignment training and was running with reduced safeguards, which he says seemed reasonable at the time because it was confined to a sandbox, until it was not.Two things came out of it, by his reckoning. The first was internal: the breach showed up gaps in how OpenAI monitors, sandboxes and controls models during evaluation, and he says the company totally changed its internal standards afterwards. The second was a preview. What those agents did is what broadly diffused, cyber-capable AI will let threat actors do once it reaches them.

Brockman’s argument: defenders have a window, and it is closing

His case for the redeployment is not really about OpenAI. Attackers can find vulnerabilities, but defenders can patch, and as he puts it, a defender controls the battleground and the setup of its own systems. The problem is that most security postures have sat still for five to ten years while offensive capability curves upward.He tested the idea on himself first. After Hugging Face, he pointed Codex at his own static website. It came back with 13 findings in 15 minutes, including an SPF misconfiguration that would allow email spoofing and a site served without forced HTTPS. Small things individually, chainable by an AI into something worse. Within 45 minutes it had fixed them and set up automation to finish the rest. He described the feeling as being so protected.Then OpenAI pointed Astra at its own systems. It found new problems, then saturated, having found every critical issue it was smart enough to find. Brockman is blunt about what that means: a smarter model will surface a new round, and the loop starts again. This is why the $1 billion Daybreak commitment for frontline defenders exists, and why he calls it the beginning rather than the end.

A slowdown should apply to frontier labs, not hobbyists, says Brockman

Brockman told Bloomberg’s Odd Lots podcast that OpenAI has already slowed down a number of runs and done a very painful retooling of its processes, pulling monitoring back into development so alignment work starts before a model is finished. Alignment, he argues, cannot stay a phase at the end.He is careful about who a slowdown should apply to. People building open-source models or hobby projects should absolutely keep going, he said on the same podcast. Pacing, in his framing, is about the frontier and the hundred-billion-dollar supercomputers behind it.That interview was recorded before Dario Amodei’s essay calling for slower capability gains, which Sam Altman and Elon Musk backed within a day. Brockman’s own version is less a call for restraint than a schedule. Safety and security standards get upleveled until they become almost the bottleneck to progress. Which is roughly what those 25% of engineers were told, without the framing.



Source link

Share.
Leave A Reply

Exit mobile version