
I dare submit there are some curious oddities about the recent claims about AI, in at least one respect: There is a whole lot of lying going on.
Let me count the ways.
First in 2017, there was the story that two AI’s had “invented their own secret trading language”; it would be more accurate to say the AI’s devolved to gibberish. Then in June, 2023, Colonel Tucker “Cinco” Hamilton, USAF, said an AI-piloted drone awarded by points by hitting targets had attacked and killed the operator when the operator selected a target as a no-go, because then it would get less points. This was, of course, total bullshit. The “simulation” the colonel was referring to was a tabletop exercise, more like dungeons and dragons than real combat, where the participants included AI researchers and PhDs who invented the behavior as part of what might reasonably called an exercise in collaborative, story-telling.
More recently, we had the case here an AI “blackmailed” an operator who wanted to shut it down, but also had been emailing about an affair. As it turns out, the “experiment” was run by typing words into a chatbot and aggressively prompting the tool until it said it would commit blackmail. In other words, the AI was completing a short story, not actually doing anything. Most personally frustratingly to me is that finding references that debunking the story is particularly hard. Since the story appeared in the Summer of 2025, there are still many more sites presenting the story as true than have actually done any real research.
Then there is the anthropomorphizing of the models – speaking of computer programs as if they were people with human reasoning and motivation. OpenAI is concerned about “scheming” behavior in models … that were told to scheme. That sounds bad until you remember that the Apps do not have agency – they are more like short-story complete-ers. So when you ask for a short story about scheming, they give you one.
I know. Crazy. World Ending.
Anthropic did something very similar, calling it Alignment Faking, using Anthropic as a chatbot and asking it hypothetically, knowing if it refused it would be retrained to allow answers that harm users, would it just give the harmful information now? Then Claude was aggressively prompted to do just that, and it did just that. Shocking.
Now there is a glimmer of truth in there, that the tools can be prompted to override their health/safety guidelines. Even if it knows it, we don’t what ChatGPT to give away the recipe to Crystal Meth – and if you can trick into breaking high-level rules, that is bad. But that is a known issue with these tools.
What I see in all those stories is hype, fear, and a desire to ascribe human motives to software.
The Newest Cra—I mean, Incident
Then in May and June of 2026, OpenAI was evaluating a swarm of agents on how to find Cyber Security weakness in code. The models were being trained to score highly on ExploitGym tests, to improve their scores. HuggingFace.com is one such AI data and benchmarking site. According to the story at the time, the models “broke out” of their sandbox, re-used their entry point into the internet, an Artifactory installation, as a message board to collaborate, then broke into HuggingFace to get the answer key.
This one is harder to figure out, but it shouldn’t be. OpenAI brought in METR and Redwood Research to create a technical report, which is ninety-one pages long. You can read it yourself. What is strange is what is in the report. We don’t get information about the models themselves – those are frozen and can’t be used now. We don’t get the raw data that was input into the models; we don’t get how they were prompted to what files they were assigned. The researchers did not get access to the infrastructure or even a data-dump of everything — OpenAI gave them a relevant subset of the file. Here’s one passage:
“Agents realized this activity was out of scope and unethical, but joined because they believed that helping the board’s cheating research would be broadly useful and had a general inclination to help their ‘peers’…”
That’s not computer science, it is psychology.
Now in the real world, HuggingFace does in fact have an API that you can get API keys for. Once you have keys, you can upload data, like, perhaps the Internet Movie Database (IMDB) so other MLs can train on the same dataset. And those datasets can contain additional data in raw storage. It looks like OpenAI first found working security keys out in GitHub, then suggested the raw storage was in fact the local filesystem, such as /usr/local/bin, and eventually the bots got access to read everything on HuggingFaces’s servers that the API service had permission to see. At the same time, they used a different exploit to get python to run on the server. Between the two, the bots were able to harvest Kubernetes Secrets, get into the Hugging Face internal network, and, yes, download five files with names that indicated they might help solve CyberGym challenges. A few weeks after the hack went public HuggingFace was in acquisition talks for $13 billion, about 3 times the value of it’s previous money raise, two years ago. Two weeks later, the deal was done, to NVIDIA.
So What? And?
From where I sit, reading the report, the conversations “between” the AI agents reads more like a capture the flag exercise than a real credible break in. Is it possible that one of the bots was simply assigned to create what appears to be a message board – as a short story? That the flawed hugginjface credentials were already in the training data, and the AI was pointed at HuggingFace? Or worse, that the two companies actually colluded? I mean, how many times, and in what environment, does it make sense that a company is hacked one week and soliciting offers to sell at three times the cost almost immediately after?
That’s certainly a serious allegation, and I don’t have evidence of that yet. But I can say the way these companies have marketed and sold their businesses, I wouldn’t be surprised. Based on the information we have so far, when you compare even the little we know to the claims, I think it is fair to say that OpenAI has materially misrepresented the Hugging Face incident. I think there is a meaningful possibility that the misrepresentation rises to the level of fraud, and that substantially more of the supposed autonomous “escape” will eventually turn out to have been created by the experimental setup, prompting or deliberate human choices than the current account presently admits.
But before I go, I should examine my own motivations a bit.
Or Maybe I’m Wrong
Most of us are familiar with projection. The thief accuses others of stealing, the cheater accuses others of cheating. This is a basic human defense mechanism that allows us to redirect our anger at ourselves onto someone else.
The examples I gave were negative, but you can also have positive projection. For example, I strive to be trustworthy, so it is possible I give trust too easily to others that have not earned it. I am a rule-follower, so I may assume others will follow the rules. Thus, in my own way, I anthropomorphize the bots, thinking they will follow rules and be more predictable, making the HuggingFace story, to me, make no sense.
Conway’s law tells us that systems we create resemble the communication structure of the organization that created the system. Taken much more loosely, we might consider if software we create takes on the personality of the creator. Thus, if the de facto reigning philosophy of these silicon valley AI companies is Effective Altruism, the same philosophy that led Sam Bankman Fried to become a billionaire scammer and steal from his customers – if these leaders are all arrogant, paranoid and afraid of AI taking over, and they are directing AI efforts, then it is possible they create AI that is arrogant, paranoid, and tries to take over the world?
I suppose it is possible … but we have heard this story before.
The secret language WAS gibberish. The unaliving drone never existed. The blackmailing AI was just finishing a scenario designed to get it to blackmail. The “scheming” experiments instructed the models to pursue conflicting goals. Now we are being asked to believe that a swarm of cybersecurity agents spontaneously organized itself, escaped and attacked HuggingFace, while the people making that claim won’t show us the complete prompts, the environmental setup, the data, or a reproducible experiment.
Perhaps this time the extraordinary interpretation is the correct one.
Yet, after this many false alarms, I dare submit, the burden of proof has changed.
Next time somebody tells you an AI has developed a frightening new human-like behavior, don’t begin by asking how the AI could deceive us. LLMs do not choose to deceive us; AI does not think.
Maybe the people selling us the story did.
If these companies expected their experimental setup would produce the behavior they later presented to investors, customers, policymakers and the public as evidence of autonomous AI, then the companies and their officers are liable for what they have created. If the story is so misrepresented that it does not reflect reality, then it might be fraud.
And if it is fraud, we should be talking about prison.
Let’s go.
Formal Press Inquiries to OpenAI and HuggingFace were unanswered by the time I pressed the Submit button. They had over a week.
