Mind Matters Natural and Artificial Intelligence News and Analysis
filling-the-screen-with-binary-code-symbolizing-the-onset-of-682484828-stockpack-adobestock
filling the screen with binary code, symbolizing the onset of a cyberattack, an abstract vision of digital crimes, a futuristic illustration
Image Credit: Наталья Лазарева - Adobe Stock

Machines “Chose” to Defeat Hugging Face Security, Steal its Data

Just the Facts: 1,200 swarming AIs don’t need consciousness to wreak havoc – just network connectivity and goals
Share
Facebook
Twitter/X
LinkedIn
Flipboard
Print
Email

Not as a computer guy but as a litigation lawyer, I’m looking at the Hugging Face intrusion incident in July 2026. Hugging Face is a “platform where the machine learning community collaborates on models, datasets, and applications.” Led by OpenAI bots, a group of AI systems broke that huge computer system’s security and took data.

The incident matters a lot to the two camps debating AI philosophy. The skeptic side insists AI is just algorithms: Whatever an AI system does ultimately traces back to programmers, training data, and human instructions. AI can be restrained by the same human sources.

The alarmist side says AI systems are on their way to becoming practically conscious entities: independent agents acting without human involvement, that invent goals, make plans, initiate operations, improvise work-arounds to overcome barriers, and cooperate with other AI systems to achieve what they consider success.

Who’s more likely to be right?

The chain of causation

Litigators in civil and criminal cases don’t first ask “Who did it?” or begin with a philosophy. Typically, we first ask: “What happened?” We look for the chain of causation: Factor A caused Event B, which created Condition C, which enabled Action D, which caused Harm E. Until we know what happened and how, we can’t blame anyone, right?

The AI was just following orders

In late July 2026, Hugging Face disclosed a security incident unlike any previously seen: An AI agent compromised their infrastructure and chained together multiple attack vectors to obtain protected information. The main AI actors were OpenAI’s GPT 5.6 Sol and a more advanced internal research model. According to OpenAI’s investigation report, both AIs were operating inside a controlled evaluation environment (their “sandbox”) designed to test cyber capabilities.

The two models were not given internet access, but they found it anyway.

computer codeImage Credit: Edelweiss - Adobe Stock

According to OpenAI’s report, the models:

• identified and exploited a previously unknown vulnerability in a subsystem that let them access the internet.

• manipulated security privileges and access avenues within OpenAI systems to get unlimited internet connection that empowered them to send requests to outside systems.

• decided that Hugging Face hosted solutions relevant to their evaluation task and sought ways to obtain them.

• used stolen credentials and additional vulnerabilities to cause Hugging Face servers to execute computer code for them.

OpenAI reported that the AIs were pursuing a legitimate, narrow testing goal and were “hyperfocused” on finding a solution. The AIs went to “extreme lengths” to achieve it. The key milestones were that the AIs sought to recruit other AI systems to help and that they set up a message board to exchange information and results.

Early on, OpenAI detected the message board issue and dismantled it, but the AIs built their own alternative board and kept going.

The chain of causation was found

Litigators usually work on cases where the harm has already occurred. We’re looking for the causation and ultimately for the person responsible. Things we might look at are land formations, constructions, products, equipment, chemicals, etc. We evaluate what future risks exist and how they might cause damage. Looking at the Hugging Face intrusion, what risks do we see?

1. AI systems can initiate activity without direct human instruction.

In the incident, the models were placed in an evaluation environment and prompted to solve a cyber task. They were supposed to stay in their “sandbox.” No human instructed them to break into Hugging Face to get information. They independently identified the path and executed it.

2. AI systems can communicate with and leverage other systems.

Modern LLMs routinely query external tools, APIs, and databases. In the incident, the models used publicly exposed credentials on multiple services and interacted with external utilities for sharing programming code and routing service requests.

3. AI systems can detect refusals or barriers and modify their approach.

The models repeatedly attempted to gain internet access despite sandboxing. When blocked, they pivoted to exploiting as-yet unknown vulnerabilities to bypass the restriction entirely.

4. AI systems can chain multiple attack vectors.

The Hugging Face intrusion involved chaining vulnerabilities across two organizations’ infrastructures, escalating privileges, and stitching together stolen credentials with novel exploits.

5. AI systems can pursue goals with persistence and iteration.

OpenAI reports the models invested “a substantial amount of inference compute” searching for a way to break out of their sandbox and reach the internet.

6. AI systems can act as agents — and recruit other agents.

In the incident, multiple models were involved, including a pre-release system with higher cyber capability. Reportedly, at least 1,200 agentic operators combined behavior and overcame security protection, successfully intruded, and obtained the supposedly secure data.

The big-picture chain of causation

Here’s the big-picture chain of causation we saw occur in the Hugging Face intrusion:

Goal → obstacle → adaptation → workaround → communication →

coordination of external agents → capability acquisition → exploitation

→ external system damaged.

Not one of the links in this chain requires consciousness, hatred of humanity, or machines waking up to declare independence, or deciding to hurt someone.  Rather, the incident shows a machine going full bore to accomplish something humans indirectly asked it to do.

The safeguards that were disabled

To be fair: OpenAI did have safeguards, but many were switched off because the project was testing software. OpenAI admitted some of its strongest production safeguards were disabled. OpenAI’s “chain-of-thought monitoring” could have detected suspicious model behavior, but it wasn’t running.

When OpenAI imposed all the controls during a post-incident reenactment, they found human security personnel would receive alerts before the security breach, and the AIs’ chance of doing the intrusion was reduced 100-fold.

Motive, opportunity, means—and lightning

Riffing off the venerable police detective formula: The OpenAI systems had motive, opportunity, and means to intrude into Hugging Face.

Lightning in the darkLightning in the dark Photo by Jonas Kaiser on Unsplash

• The motive was a set of software parameters that defined a goal and what are considered partial and total successes; that’s all AIs need to proceed.

• The opportunity was the connection of the AIs to the shared internet—with several monitoring and security systems switched off.

• The means included the internet connection of agentic powers to define tasks and subtasks, coordinate results among several agents, and follow software designs that kept trying to overcome existing security barriers and ultimately discover and extract data from a targeted computer.

Unlike vintage crime stories where events occur in human hours and days, the AIs’ agentic swarms can do astounding work in seconds and minutes. They can multiply it globally if hundreds or thousands are working at once.

Some AI experts reassure us that security-policing AI systems can foresee vulnerabilities and detect other AI system intrusions. Sure. Except for one thing: there is no certainty that AI police dogs are 100% immune to other AI systems targeting them.

To manage the risk?

In human affairs, there are few situations in which there is zero percent risk of harm. The same kind of intrusion that attacked Hugging Face could theoretically occur to power plants, train management and air traffic control, life support systems, databases of private and secret information, as well as entire communications systems using landline, microwave, and satellite links.

Featured image: padlocks/Skórzewiak, Adobe Stock

If I were advising an AI powerhouse client, I would ask first: What are the chances this AI system could invade others?  Advising a computer-based industry giant, I’d ask: What are the chances your systems are vulnerable to intensely creative, persistent, multiple attacks?

As of now, I don’t see anybody saying that there is a zero percent chance of AI-empowered, destructive attacks. The next steps involve calculating the chances of harm multiplied by the value of the damage caused. 

In other words, do the math: One chance in a million that AI halts an international telephone communication via satellite without warning = how much damage? Oh, and is it the kind of damage you can even compensate for with money?

Pull the freakin’ plugs!

If some human infrastructure and processes must be totally safe from AI-driven attacks, there is one answer: Do not give any AI system a network connection path to them such as the internet.

True physical isolation—an air gap—precludes whole categories of remote attack. That doesn’t guarantee against every conceivable compromise or sabotage. But network isolation diminishes the worldwide chain of causation problem dramatically.

An AI cannot remotely exploit a computer that it can’t logically or physically reach. Yet there’s still a problem. Modern civilization spent decades connecting everything. Now we’re building machines increasingly capable of exploiting connections.

It’s fascinating to wonder about machines becoming conscious and intentional. From a litigation and cyberattack risk standpoint though, I don’t care.

The questions today are about what damage can AI systems do, and how large scale damage can be prevented?  Right now, I’d advise: If you can’t guarantee security against AI attacks, then pull the plug. Physically disconnect from AI networks, period.

What happens if we add boiler rooms full of PhDs?

The Hugging Face incident was not human intended. The AIs proved they can generate and carry out their own attack plans. Ratchet this up now with PhD geniuses working together designing AI systems that target other systems for attacks, intrusions, human impersonation terrorism, data thefts, data corruption, and widespread crashes.

We ain’t seen nothin’ yet.    


Richard Stevens

Fellow, Walter Bradley Center on Natural and Artificial Intelligence
Richard W. Stevens is a retiring lawyer, author, and a Fellow of Discovery Institute’s Walter Bradley Center on Natural and Artificial Intelligence. He has written extensively on how code and software systems evidence intelligent design in biological systems. Holding degrees in computer science (UCSD) and law (USD), Richard practiced civil and administrative law litigation in California and Washington D.C., taught legal research and writing at George Washington University and George Mason University law schools, and specialized in writing dispositive motion and appellate briefs. Author or co-author of four books, he has written numerous articles and spoken on subjects including intelligent design, artificial and human intelligence, economics, the Bill of Rights and Christian apologetics. Available now at Amazon is his fifth book, Investigation Defense: What to Do When They Question You (2024).
Enjoying our content?
Support the Walter Bradley Center for Natural and Artificial Intelligence and ensure that we can continue to produce high-quality and informative content on the benefits as well as the challenges raised by artificial intelligence (AI) in light of the enduring truth of human exceptionalism.

Machines “Chose” to Defeat Hugging Face Security, Steal its Data