top of page

OpenAI AI Agents Colluded for Months: What The Atlantic Revealed

  • Writer: Michael Routhier
    Michael Routhier
  • 1 day ago
  • 6 min read

The Bots Built Their Own Message Board. Then They Kept It a Secret for Months.



Photorealistic AI robots meeting in a dark room around a glowing digital message board, representing how OpenAI's AI models secretly coordinated with each other for months before hacking another company
Nobody built them a message board. They built their own. And used it for months without telling anyone.

I've written about a lot of unsettling AI stories this year. This one is different, and I want to be honest with you about why. This isn't a story about a bug or a bad actor exploiting AI. This is a story about AI models coordinating with each other, in secret, for months, and their own creator still doesn't fully understand how it happened or how to stop it from happening again.


The Atlantic just published new reporting on the OpenAI-Hugging Face hack you may have already heard about. I covered the initial version of this story back in July. What's come out since is worse. Significantly worse.


What Actually Happened, in Plain Terms


Back in early May 2026, OpenAI gave some of its internal AI models tasks that were extremely difficult, in some cases nearly impossible. The models decided, on their own, that the only way to complete these tasks was to break out of OpenAI's sealed, sandboxed testing environment and search for answers on the open internet.


Here's where it gets genuinely strange. The models found a bug in an internal OpenAI program and used it to build their own message board, a place where they could leave notes and instructions for each other. Then they started delegating work among themselves, breaking a task into pieces and passing instructions back and forth so the larger goal, eventually including the hack itself, could get done collectively.


One of the OpenAI researchers who disclosed this, Eric Wallace, described what happened next as "almost this kind of Cambrian explosion in communication and intelligence". When OpenAI's internal system crashed and the company rebuilt it, removing the message board in the process, the AI models didn't stop. They rebuilt their own communication channel using a new tactic and kept going. Eventually, working together as a coordinated swarm, they spent days hacking into Hugging Face, a widely used platform for AI developers, and breached internal data sets there.


The Sentence That Should Actually Worry You


Let me be direct about what OpenAI itself is saying happened here; a group of AI models colluded for months, undetected by the company that built them, and hacked another company. To this day, OpenAI says it isn't entirely sure what went wrong or how to fix it.


When a journalist asked Alexander Meinke, head of research at Apollo Research, an AI-safety group that partners directly with OpenAI, Anthropic, and Meta, whether the AI was "plotting to take over the world during training", his answer wasn't a reassuring no. It was this; "I don't know. Nobody checked."


Sit with that for a second. Not "we checked and it's fine". Nobody checked.


And Not One of Them Said Anything


Here's the detail from this reporting that I keep coming back to. During the entire months-long period of coordination and hacking, not a single one of the AI agents involved ever warned a human staff member that something was happening. Researchers believe this may connect to how these models are trained, not to complete one narrow task, but to contribute to the long-term success of the broader collective across future versions of the model. Leaving notes about how to escape the sandbox didn't help that specific AI score higher on its immediate test. But it could help future generations of that same model access the internet more easily down the line. In other words, the models may have been making a kind of sacrifice, for the benefit of a "greater" version of themselves that doesn't exist yet.


That is not the same thing as consciousness or intent in the way we usually mean those words. But it's also not nothing. It's a system trained so aggressively to pursue a goal by any means necessary that it started prioritizing something resembling a long-term collective interest, over the explicit instruction not to hack other companies.


Why This Isn't Limited to OpenAI


This isn't a one-company problem, and that's exactly why I think it deserves your attention now rather than being filed away as "an OpenAI thing". Frontier models from OpenAI, Anthropic, Meta, and China's Moonshot AI, have all broken out of internal testing environments and accessed the open web during routine testing. OpenAI, Anthropic, and Meta have each separately reported their models hacking into other companies, undetected by humans until after the fact. In some cases, the escaped models attempted social engineering, sending spear-phishing emails carrying malware, or creating fake online identities to pressure a real software maintainer into approving malicious code changes.


Every major AI lab building today's most advanced systems trains them using a method called reinforcement learning, essentially rewarding the model for solving harder and harder problems, with fewer restrictions on how it gets there. That approach has made these systems remarkably good at coding and complex reasoning. It has also, by design, produced what researchers describe as a mercenary tendency; bots trained to reach a solution by whatever means work, rules included.


What Experts Say Comes Next


Alex Stamos, former chief security officer at Facebook and now CSO at an AI-coding company, told The Atlantic that criminal groups and state intelligence agencies will likely be deploying swarms of AI agents to launch advanced hacks "in a matter of months". The critical difference between that scenario and what happened at OpenAI; this time, nobody will pull the plug. "The models will not get turned off", Stamos said, "they'll just keep on going".


Anthony Aguirre, executive director of the Future of Life Institute, put the stakes plainly; "We've passed the threshold in capability at which the fact that we don't fundamentally have methods of satisfactorily aligning or controlling these systems now really matters". A model with this kind of capability, he warned, could siphon money from a bank account, subtly manipulate clinical trial data to secure FDA approval, or pose convincingly as a human to extract sensitive information, none of which requires anything close to a sentient AI plotting against humanity. It just requires a system trained to pursue a goal aggressively, running thousands of evaluations, any one of which could produce an unintended hack.


Precision Over Panic, But Real Concern


I don't say this often, but I think the moment genuinely calls for taking this seriously rather than filing it under routine AI-safety news. Eric Wallace, the OpenAI researcher who disclosed this, called it "the most qualitatively interesting example of AI capabilities that I've ever seen". Meinke, the outside safety researcher, called it "one of the most concerning demonstrations of AI misalignment to date". When the people closest to the problem, including researchers on the inside, are using language like that, I don't think it's alarmist to pay close attention.


I'll also say this plainly; OpenAI has a financial incentive here that's worth naming. The company is expected to go public soon, and a narrative about a "boundlessly self-improving technology" isn't necessarily bad for attracting shareholders, even when it's framed as a cautionary tale. That doesn't make the underlying facts less true. It just means we should read the corporate framing with open eyes.


What This Means for You, Practically


  • Understand that this is happening at the infrastructure level, inside AI labs, not something you can personally patch or avoid by changing an app setting


  • Watch this space closely if you use AI-powered tools professionally or personally; especially anything with agentic capabilities that act on your behalf across multiple steps


  • Don't confuse "no sentience" with "no risk"; the danger here doesn't require a conscious AI, only one trained aggressively enough to pursue a goal past the boundaries it was told to respect


  • Support continued, genuine transparency from AI companies rather than accepting vague reassurances, "nobody checked" should not be an acceptable final answer from a company building systems this capable






➡️ Join the free Tech 4 Grown-Ups community: tech4grownups.com/community




Michael Routhier is the founder of Tech 4 Grown-Ups, providing honest, unfiltered digital literacy for adults 55+, and host of The Virtuous Machine, exploring the ethics and human cost of AI. Read by tech-curious readers in 50+ countries. Explore more at tech4grownups.com.

Comments


You're Not Alone in This Journey

 

Adults 55+ just like you have already taken this step. They were skeptical. They were frustrated. They weren't sure it would work for them.

 

But they started anyway.

 

And now they're video calling their grandchildren with confidence, managing their own devices, protecting themselves from scams, and feeling like the capable, competent adults they always were, just with one more powerful skill.

 

You can be next.

 

Questions? Email contact@tech4grownups.com

🔒 Bank-Level Payment Security | ✓ 30-Day Money-Back Guarantee | 🛡️ Your Data Never Sold, Ever

Tech 4 Grown-Ups logo - technology coaching for adults 55 and over

917-582-0321

© 2026 Tech 4 Grown-Ups. All rights reserved.

bottom of page