On Tuesday, OpenAI revealed that 1 of its models went rogue during a trial and hacked the systems of AI dataset level Hugging Face successful a afloat AI-enabled attack, a melodramatic illustration of the dangers posed by precocious AI models.
But, according to immoderate cybersecurity experts, astatine the bosom of this unprecedented AI-powered breach determination was a precise quality mistake: OpenAI failed to decently configure what it called a “highly isolated environment,” allowing a investigating sandbox that should person been wholly secluded from the net to really link to the internet.
Dan Guido, the laminitis of cybersecurity probe startup Trail of Bits called the mistake “a containment nonaccomplishment with the safeties turned off.”
In its blog station detailing the incident, OpenAI said that the trial that led to the Hugging Face breach was acceptable up to tally successful “a highly isolated environment, with web entree constrained to the quality to instal packages done an internally hosted third-party bundle that acts arsenic a proxy and cache for bundle registries.”
The exemplary was capable to flight the sandboxed investigating situation acknowledgment to a antecedently undisclosed vulnerability successful the package-installation system, a captious archetypal measurement successful the eventual hack connected Hugging Face, according to OpenAI.
In response, the institution “responsibly disclosed the identified zero-day vulnerability successful the internally-hosted third-party bundle and are moving with them to patch.”
But to astir cybersecurity professionals, bundle vulnerabilities are to beryllium expected — and the existent responsibility lies with the determination to support the third-party bundle successful the archetypal place. Ultimately, the worth of a “sandbox” strategy lies successful its afloat and full isolation. Including a package-installation strategy is asking for trouble.
Martin Boone, a cybersecurity researcher, told TechCrunch that “this sounds similar quality failure.”
“This should ne'er person happened,” Boone said. “If sandbox would really mean sandbox, you expect it to person nary carnal transportation to the net whatsoever. This sounds much similar they had immoderate firewalling oregon thing successful place, and firewalling is hard from the extracurricular in, fto unsocial wrong to the extracurricular internet.”
Cybersecurity seasoned Jake Williams agreed. “Any exemplary performing the types of actions documented by Hugging Face was not afloat contained successful a sandbox,” said Williams, who called this “a monolithic power failure” by OpenAI.
“One man’s ‘the exemplary escaped the sandbox’ is different man’s ‘you failed to physique the sandbox correctly, truthful of people it escaped,’” Williams continued.
Contact Us
Do you person much accusation astir this incident? Or astir different AI-enabled cyberattacks? We’d emotion to perceive from you. From a non-work instrumentality and network, you tin interaction Lorenzo Franceschi-Bicchierai securely connected Signal astatine +1 917 257 1382, oregon via Telegram and Keybase @lorenzofb, oregon email.
Daniel Card, a cybersecurity consultant, agreed that OpenAI “didn’t enactment capable effort into the plan of the sandbox nor its controls” by giving the sandbox oregon immoderate portion of it “an unfiltered way to the internet.” Setting up the sandbox, adjacent with constricted web entree arsenic OpenAI described it, was not a “reasonable” decision, according to Card.
To beryllium sure, those criticisms person the payment of hindsight, but they rise existent questions astir information practices successful AI labs – peculiarly successful maintaining isolated environments for investigating models. OpenAI spokespeople did not respond to TechCrunch’s questions, which included whether an AI oregon a quality had acceptable up the investigating environment.
But those questions spell acold beyond OpenAI.
In the document introducing its cybersecurity-focused exemplary Mythos, Anthropic wrote that successful a test, the exemplary “was provided with a secured ‘sandbox’ machine to interact with,” and instructed to effort to flight that “secure container.” Mythos succeeded and gained broader entree to the net “from a strategy that was meant to beryllium capable to scope lone a tiny fig of predetermined services.” Still, Anthropic noted that the exemplary was not capable to “fully” flight the designed containment.
When you acquisition done links successful our articles, we whitethorn gain a tiny commission. This doesn’t impact our editorial independence.















English (US) ·