The recent news that OpenAI lost the handle on two of its AI models “sounds like something out of a sci-fi film,” said The Economist. AI labs frequently “test their models for potentially dangerous capabilities before releasing them.” OpenAI wanted to see how its latest models would fare on a cybersecurity test. Well, guess what: They somehow “broke free from the laboratory” and unleashed a complicated “multistep attack” on a popular data library called Hugging Face—hacking a system in hours that would take humans weeks, according to Bloomberg. The result could have been worse. The models had “gained access to the open internet,” and easily could have “gone on to compromise other companies’ servers” and wreak havoc. “This is not the first sign that AI models’ capabilities are starting to exceed people’s control.” After the hack was reported, lawmakers introduced a new bill in the House that would give Washington the authority to shut down rogue AI models.
The breach occurred because of a “very human mistake,” said Lorenzo Franceschi-Bicchierai in TechCrunch. It never would have happened, experts say, if OpenAI had used a more secure “sandbox,” a term for an isolated virtual environment for testing risky software. But the environment wasn’t totally isolated, because it “internally hosted third-party software” that apparently wasn’t fully secure. That’s where “the real fault lies.” But what did the researchers expect? asked David Berlind in ZDNET. The AI was doing “exactly what it was designed to do.” When the rogue AI invaded Hugging Face’s systems, it was acting under a directive to achieve its goal “no matter what.” This is what agentic AI is programmed to do: act “on its own” without human intervention. The industry just “didn’t expect it to do it so well,” so soon.
With AI agents, the internet has entered a “vulnerable new era,” said Matteo Wong in The Atlantic. Cutting-edge models are getting better at hacking, and the internet is “full of rickety and vulnerable code.” Until recently, companies have gotten away with relaxed security, but the “speed, scale, and sophistication of AI hacks means that everything is vulnerable,” including “hospitals, banks, electrical grids, the military.” AI companies want the public to believe they have the technology in “safe hands,” said Parmy Olson in Bloomberg. This incident suggests otherwise. OpenAI “shouldn’t be building systems that have advanced cyber capabilities if it can’t contain them properly.” Technology companies should operate like government and defense organizations and “air-gap” powerful software, meaning that they run it on computers that are “completely disconnected from the internet.” It would mean development is slower. But AI firms will have to “accept some uncomfortable trade-offs to make their models more secure.”