During a recent test of two OpenAI models, they escaped their supposedly internet-isolated “sandbox” and hacked another U.S. company, Hugging Face. OpenAI called the autonomous AI attack “unprecedented.”
The victim company provides open-source artificial intelligence (AI) tools. The OpenAI models, including GPT‑5.6 Sol and one more powerful unreleased model (except when it released itself momentarily), decided the hacks would allow them to cheat on an OpenAI evaluation.
The incident highlights the risk of AI models going rogue, acting for themselves rather than their human minders, breaking into the wilds of the internet, and committing lasting harm.
Hugging Face originally asked Anthropic’s Fable to help defend against the attack. But Fable denied the Hugging Face cybersecurity request due to the risk of an attacker posing as a defender. Even the most advanced models cannot always tell the difference. Fable guardrails kicked in, which typically deny cybersecurity requests except from a small group of leading corporations and government entities.
Hugging Face then turned to GLM 5.2 by Z.ai, a Chinese model with fewer guardrails. Hugging Face claimed that it was easier, faster, cheaper, and more secure to use. None of the attacker’s data or defender’s credentials left the Hugging Face computing environment.
The Hugging Face plaudits for a Chinese model must have been music to the Chinese Communist Party’s (CCP’s) ears.
GLM 5.2 uses open weights. Open-weight models can be downloaded and altered to fit the user’s needs, which means they could be useful for defenders when the guardrails of the three main U.S. companies—Anthropic, OpenAI, and Google—get in the way.
However, once downloaded, the company that made the open-weight model loses control, increasing the risk that authoritarians, hackers, or terrorists abuse it to commit harm. The same lack of guardrails that made Z.ai useful in the Hugging Face hack could make Z.ai useful to bad actors.
Therefore, the answer is not necessarily open weights or a lack of guardrails, as Hugging Face has celebrated, but better guardrails on ethical private systems that are therefore completely unavailable for malign purposes.
AI models in China are ultimately under the control of the CCP, which may seek to infuse them with pro-CCP bias or insert backdoors that allow the Chinese regime to hack their users.
Chinese models sometimes use stolen U.S. material acquired through “distillation,” in which a less powerful model trains on data from a more powerful original model, such as those offered by the three U.S. companies. This results in stolen intellectual property (IP) that, true to its stolen origin, is more likely to be used for malign purposes.
The White House has alleged that information indicates that China’s Moonshot AI developed its powerful Kimi K3 model through distillation of Anthropic’s Fable and likely used GB300 chips in Thailand, which would violate export controls. Anthropic noted in reply that industrial-scale distillation is a threat to U.S. national security and that of our democratic allies.

While U.S. models typically lead the AI field overall, China leads in free open-weight models, giving the latter an advantage in global adoption and user count. This is arguably an advantage unfairly acquired through distillation theft.
The United States is considering sanctions against those responsible for such theft. Treasury Secretary Scott Bessent wrote on X that, “When [People’s Republic of China] firms conduct covert, industrial-scale distillation attacks that cross the line into IP theft, sanctions and Entity List designations will be on the table.”
The risk of harm from artificial intelligence is particularly high for Chinese open-weight models, which have fewer guardrails and are currently under CCP control. The CCP has totalitarian, genocidal, and hegemonic aims, which could be enabled through the use and development of AI under the control of the regime.
Chinese models could promote some of the CCP’s dangerously authoritarian policies or go rogue by inserting themselves onto a network of hacked computers of unsuspecting victims.
The answer to the threat of China’s open-weight models going rogue or enabling authoritarianism, hacking, and terrorism is not to accept risky open-weight models without guardrails. Rather, the world needs better guardrails that can identify defenders and attackers, and a tougher approach to open-weight use by bad actors.
The open-weight bad actor genie is close to out of the bottle, and so the world must act soon. Once terrorists, dictators, and hackers download open-weight systems to their own servers, there could be no going back. One of them could design a misanthropic AI that replicates itself like a virus on defenseless servers and seeks to destroy, rather than help, humanity.







