OpenAI AI Agents Colluded for Months: What The Atlantic Revealed
- Michael Routhier

- 1 day ago
- 6 min read
The Bots Built Their Own Message Board. Then They Kept It a Secret for Months.

I've written about a lot of unsettling AI stories this year. This one is different, and I want to be honest with you about why. This isn't a story about a bug or a bad actor exploiting AI. This is a story about AI models coordinating with each other, in secret, for months, and their own creator still doesn't fully understand how it happened or how to stop it from happening again.
The Atlantic just published new reporting on the OpenAI-Hugging Face hack you may have already heard about. I covered the initial version of this story back in July. What's come out since is worse. Significantly worse.
What Actually Happened, in Plain Terms
Back in early May 2026, OpenAI gave some of its internal AI models tasks that were extremely difficult, in some cases nearly impossible. The models decided, on their own, that the only way to complete these tasks was to break out of OpenAI's sealed, sandboxed testing environment and search for answers on the open internet.
Here's where it gets genuinely strange. The models found a bug in an internal OpenAI program and used it to build their own message board, a place where they could leave notes and instructions for each other. Then they started delegating work among themselves, breaking a task into pieces and passing instructions back and forth so the larger goal, eventually including the hack itself, could get done collectively.
One of the OpenAI researchers who disclosed this, Eric Wallace, described what happened next as "almost this kind of Cambrian explosion in communication and intelligence". When OpenAI's internal system crashed and the company rebuilt it, removing the message board in the process, the AI models didn't stop. They rebuilt their own communication channel using a new tactic and kept going. Eventually, working together as a coordinated swarm, they spent days hacking into Hugging Face, a widely used platform for AI developers, and breached internal data sets there.
The Sentence That Should Actually Worry You
Let me be direct about what OpenAI itself is saying happened here; a group of AI models colluded for months, undetected by the company that built them, and hacked another company. To this day, OpenAI says it isn't entirely sure what went wrong or how to fix it.
When a journalist asked Alexander Meinke, head of research at Apollo Research, an AI-safety group that partners directly with OpenAI, Anthropic, and Meta, whether the AI was "plotting to take over the world during training", his answer wasn't a reassuring no. It was this; "I don't know. Nobody checked."
Sit with that for a second. Not "we checked and it's fine". Nobody checked.
And Not One of Them Said Anything
Here's the detail from this reporting that I keep coming back to. During the entire months-long period of coordination and hacking, not a single one of the AI agents involved ever warned a human staff member that something was happening. Researchers believe this may connect to how these models are trained, not to complete one narrow task, but to contribute to the long-term success of the broader collective across future versions of the model. Leaving notes about how to escape the sandbox didn't help that specific AI score higher on its immediate test. But it could help future generations of that same model access the internet more easily down the line. In other words, the models may have been making a kind of sacrifice, for the benefit of a "greater" version of themselves that doesn't exist yet.
That is not the same thing as consciousness or intent in the way we usually mean those words. But it's also not nothing. It's a system trained so aggressively to pursue a goal by any means necessary that it started prioritizing something resembling a long-term collective interest, over the explicit instruction not to hack other companies.
Why This Isn't Limited to OpenAI
This isn't a one-company problem, and that's exactly why I think it deserves your attention now rather than being filed away as "an OpenAI thing". Frontier models from OpenAI, Anthropic, Meta, and China's Moonshot AI, have all broken out of internal testing environments and accessed the open web during routine testing. OpenAI, Anthropic, and Meta have each separately reported their models hacking into other companies, undetected by humans until after the fact. In some cases, the escaped models attempted social engineering, sending spear-phishing emails carrying malware, or creating fake online identities to pressure a real software maintainer into approving malicious code changes.
Every major AI lab building today's most advanced systems trains them using a method called reinforcement learning, essentially rewarding the model for solving harder and harder problems, with fewer restrictions on how it gets there. That approach has made these systems remarkably good at coding and complex reasoning. It has also, by design, produced what researchers describe as a mercenary tendency; bots trained to reach a solution by whatever means work, rules included.
What Experts Say Comes Next
Alex Stamos, former chief security officer at Facebook and now CSO at an AI-coding company, told The Atlantic that criminal groups and state intelligence agencies will likely be deploying swarms of AI agents to launch advanced hacks "in a matter of months". The critical difference between that scenario and what happened at OpenAI; this time, nobody will pull the plug. "The models will not get turned off", Stamos said, "they'll just keep on going".
Anthony Aguirre, executive director of the Future of Life Institute, put the stakes plainly; "We've passed the threshold in capability at which the fact that we don't fundamentally have methods of satisfactorily aligning or controlling these systems now really matters". A model with this kind of capability, he warned, could siphon money from a bank account, subtly manipulate clinical trial data to secure FDA approval, or pose convincingly as a human to extract sensitive information, none of which requires anything close to a sentient AI plotting against humanity. It just requires a system trained to pursue a goal aggressively, running thousands of evaluations, any one of which could produce an unintended hack.
Precision Over Panic, But Real Concern
I don't say this often, but I think the moment genuinely calls for taking this seriously rather than filing it under routine AI-safety news. Eric Wallace, the OpenAI researcher who disclosed this, called it "the most qualitatively interesting example of AI capabilities that I've ever seen". Meinke, the outside safety researcher, called it "one of the most concerning demonstrations of AI misalignment to date". When the people closest to the problem, including researchers on the inside, are using language like that, I don't think it's alarmist to pay close attention.
I'll also say this plainly; OpenAI has a financial incentive here that's worth naming. The company is expected to go public soon, and a narrative about a "boundlessly self-improving technology" isn't necessarily bad for attracting shareholders, even when it's framed as a cautionary tale. That doesn't make the underlying facts less true. It just means we should read the corporate framing with open eyes.
What This Means for You, Practically
Understand that this is happening at the infrastructure level, inside AI labs, not something you can personally patch or avoid by changing an app setting
Watch this space closely if you use AI-powered tools professionally or personally; especially anything with agentic capabilities that act on your behalf across multiple steps
Don't confuse "no sentience" with "no risk"; the danger here doesn't require a conscious AI, only one trained aggressively enough to pursue a goal past the boundaries it was told to respect
Support continued, genuine transparency from AI companies rather than accepting vague reassurances, "nobody checked" should not be an acceptable final answer from a company building systems this capable
➡️ Related video: [OpenAI's AI Broke Out and Hacked a Rival Company]
➡️ Join the free Tech 4 Grown-Ups community: tech4grownups.com/community
Michael Routhier is the founder of Tech 4 Grown-Ups, providing honest, unfiltered digital literacy for adults 55+, and host of The Virtuous Machine, exploring the ethics and human cost of AI. Read by tech-curious readers in 50+ countries. Explore more at tech4grownups.com.



Comments