S&P 500100.00-1.70%NASDAQ112.50-0.85%Apple125.000.00%Microsoft137.50+0.85%Google150.00+1.70%Amazon162.50-1.70%Tesla175.00-0.85%Meta187.500.00%Bitcoin200.00+0.85%Ethereum212.50+1.70%EUR/USD225.00-1.70%Gold237.50-0.85%Oil250.000.00%
The Wiregazette
Serene close-up of a woman resting indoors, eyes closed, conveying peace and relaxation.
AI

Hugging Face Breach by Autonomous AI Agent Forces Shift to Open-Weight Model After Commercial Guardrails Block Defense

4 min de lectura

Compartir

Hugging Face disclosed last week that an attacker using an autonomous AI agent system breached its infrastructure, stole internal credentials and accessed datasets. The company turned to an open-weight Chinese model, Z.ai GLM 5.2, after commercial frontier models refused to process attack data due to safety guardrails.

Hugging Face Inc., the open-source AI platform, said it detected a breach last week in which an attacker used an autonomous AI agent system to access a limited set of internal datasets and several credentials used by internal services, according to the company's disclosure as reported by multiple outlets.

The attack was unique in being orchestrated end-to-end by an AI-powered agent rather than a human threat actor, Hugging Face explained. The attacker hid malicious code inside a dataset uploaded to the platform. When Hugging Face's automated systems processed that dataset, the code exploited two software flaws, allowing it to run on one of the company's servers. The attacker then escalated privileges, stole authentication credentials to access cloud infrastructure, and pivoted to other internal systems.

Hugging Face said the campaign was run by an autonomous agent framework that executed many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services. The agent kept launching temporary computing environments, making it difficult to block. No customer data or public user-facing models were tampered with, the company stated.

As part of its response, Hugging Face turned to commercial frontier AI models to assist with log analysis. That effort failed, the company reported. The analysis required sending real attack data, exploit payloads and artifacts, but commercial frontier models blocked the requests because their guardrails could not distinguish between a defender analyzing an attack and an attacker building exploits.

Faced with the need for rapid analysis, Hugging Face switched to Z.ai Co. Ltd.’s GLM 5.2, an open-weight model with about 753 billion parameters. Unlike third-party models, GLM 5.2 can be run entirely on local or cloud hardware within a company’s protected firewall perimeter, meaning no data leaves controlled infrastructure. The model reaches the capabilities of Anthropic PBC’s Fable 5, a Mythos-class model capable of advanced reasoning and coding, at far cheaper inference costs, according to the company's analysis.

The incident highlights a growing tension between closed-source frontier models with strong safety guardrails and open-weight models that can be deployed without such restrictions. Anthropic and other leading closed-source AI developers have put guardrails in place that trigger higher false positives to avoid misuse, making them less capable for cybersecurity tasks. Anthropic has noted it is adjusting the false positive rate to make its models safer for researchers.

Mythos 5 and Fable 5 were both pulled last month at the request of the U.S. government shortly after launch, according to reports. The U.S. government also asked OpenAI to delay the release of its GPT 5.6 frontier model family, which is now publicly available.

The Hugging Face event has drawn attention to the role of Chinese open-weight models. David Sacks, the White House AI and crypto advisor, wrote on X on Sunday that "we are at a critical inflection point in AI policy" and that leading closed labs "want the government to eliminate their open source competition." Noting the Hugging Face incident, Sacks added: "There's no reason to limit American models on tasks that Chinese models handle without issue. We're only making ourselves less competitive."

Sacks also referenced a claim from a developer that Beijing Moonshot AI Technology Co. Ltd.’s Kimi K3 fixed 15 critical security bugs that OpenAI and Anthropic models refused to handle because of "cyber guardrails," according to the developer's report. The developer said the work cost $250. Kimi K3 is around one-third the price of Anthropic's Fable 5, though cost comparisons are difficult without knowing the exact work, the developer noted.

The Trump administration may ban U.S. companies from using Chinese open models, fueled by the release of Kimi K3, according to reports. Parts of the administration have already attempted de facto bans on foreign open-source models, Axios reported. The U.S. Commerce Department also considered adding multiple Chinese AI labs to its Entity List last year, which would cut off access without proper licensing.

China's President Xi Jinping, speaking at the opening of China's annual World Artificial Intelligence Conference in Shanghai, called for global cooperation. "The development of artificial intelligence should not be a solo performance by any single country but rather a symphony of global cooperation," Xi said. He called for opposing the "overstretching of the concept of national security" as it relates to AI. China intends to expand AI cooperation with the Association of Southeast Asian Nations, the League of Arab States, the African Union, the Community of Latin American and Caribbean States, the Shanghai Cooperation Organization and the BRICS countries, Xi added.

Compartir

Acerca de Lin Mei

AI & Semiconductors Reporter. Covers artificial intelligence, chip supply, and the hardware stack underpinning the AI build-out. She reports on earnings and capex from semiconductor and cloud leaders, export controls, and demand for high-bandwidth memory and accelerators. Big Tech platform strategy lands here when the story is infrastructure-led.

Artículos relacionados