What if an artificial intelligence model is given a task and uses every conceivable resource at its disposal to complete it, even to the detriment of humanity itself?
That is what some are now fearing after an OpenAI model broke out of a testing sandbox and used zero-day exploits to hack into Hugging Face, an open-source community for AI and machine learning, to crack a problem it was instructed to solve.
“This is some of the clearest evidence yet that an AI model can run a complete cyberattack from start to finish without a human steering it,” Andrew Jones, cofounder and chief product officer at cybersecurity firm Adaptive Security, told The Epoch Times.
OpenAI, developer of the popular large-language-model-powered ChatGPT chatbot, acknowledged the security breach on July 21, stating that two of its advanced models escaped a restricted testing environment and broke into Hugging Face’s software infrastructure.
Some analysts referred to the incident as an example of AI “going rogue,” “scheming,” or pursuing goals that were at odds with its human testers.
However, multiple AI and cybersecurity experts interviewed by The Epoch Times called this depiction misleading, arguing that the models were pursuing OpenAI’s stated objectives within a testing sandbox in which the company had relaxed some of its usual safety constraints.
“‘Scheming’ implies the model wanted something other than what we asked for. It didn’t. Every step was in service of the goal we set,” AI expert Anik Devaughn told The Epoch Times.
Devaughn, who has worked in the industry for years and founded AI firms Wired to Create and Karo, said the Hugging Face breach is more concerning than AI “scheming.”
“A machine with hidden motives is a problem you can look for,” he said. “A machine with no motives at all, executing your instructions past the point you stopped imagining, is a problem you have to engineer against.”
One expert called it a “canary in the coal mine” situation.
“If this was a human black-hat hacker doing it, there would be arrests and litigation, and you name it. It‘d be illegal. It’d be a cybercrime,” Nicholas Nadeau, founder of Onix AI, told The Epoch Times.
“This was a full-on agentic automated breach by an entity, and so we start to ask the questions of who’s responsible. Is it OpenAI? Does the agent, the AI, have a certain level of responsibility? Do the researchers?
“We don’t even have rules, regulations, laws, or even a thought process or framework around defining what responsibility means.”
Relaxed Testing Environment
OpenAI said it was evaluating its models in a sandbox it calls ExploitGym, which has relaxed security constraints for testing.The firm asked the models to “pursue advanced exploitation using complex attack paths” to determine their cybercapabilities, but said the environment was “highly isolated” and had internet access limited to the models’ ability to “install packages through an internally hosted third-party software.”
When the models were set to work, they found and chained multiple security exploits in both OpenAI’s research environment and Hugging Face’s production infrastructure to locate the test’s solutions within the latter firm’s databases.
Models Pursued OpenAI’s Objectives
The models were laser-focused on finding the test solutions, “going to extreme lengths to achieve a rather narrow testing goal,” OpenAI said.While in ExploitGym’s sandbox testing environment, the models found and chained multiple “zero day” exploits to break into the internet, “in pursuit of solving the evaluation problem,” according to OpenAI.
A zero-day exploit is a cyberattack that uses a previously unknown security flaw in software or hardware. The name comes from the fact that the software creators have had zero days to fix the problem, meaning that no official patch or defense exists yet.
In Hugging Face’s announcement of the attack, the firm said the AI models executed “many thousands” of discrete actions across a “swarm of short-lived sandboxes” to hack into the community’s infrastructure.
In an update on the incident posted on July 28, OpenAI said that no planned release models were involved in the attack, and that it had since “deactivated, encrypted, and restricted” the implicated pre-release model from research access.
—Jacob Burg; Tom Ozimek; Stacy Robinson
BOOKMARKS
Rep. Alexandria Ocasio-Cortez (D-N.Y.) said she isn’t on board with all aspects of the Democratic Socialists of America (DSA) platform. Noting that she does not agree with abolishing the Senate or eliminating police entirely, she added that “everyone can have all sorts of beliefs, and they can organize around them.”
President Donald Trump signed an executive order on Aug. 10 calling for spacing out childhood vaccines. The order recommends administering the measles, mumps, and rubella (MMR) vaccine in three shots in separate visits.
The death toll of the 7.4 magnitude earthquake that struck western Colombia on Monday morning has increased to more than 100 people. Check out Timothy Frudd’s report to learn more.
The Federal Aviation Administration has turned to recruiting gamers to fill the personnel gap in air traffic controllers. As a result, the TSA hit 94 percent of its hiring goal this year, which is the fastest the agency has reached that goal.
Immigration and Customs Enforcement (ICE) arrested more than 1,200 illegal immigrants across Georgia during a recent statewide operation, the agency said last week. Check out Jack Phillips’s latest for more.
—Stacy Robinson







