Ominous warning as AI swarms threaten civilization
AI plans, lies, cheats, coordinates, and collaborates to achieve a mission.
“In the coming months, AI-enabled cyber attacks will become far more widespread and sophisticated as models around the world become increasingly capable. The companies and public services our communities depend on—from hospitals to water treatment plants to the infrastructure that powers the internet—are at risk.” - OpenAI in an open letter calling for a surge in global cyber defense.
Developers are pounding the table but nobody’s listening. The reality is those calling for action are the ones creating the most damage.
When you hear “AI” you might think of LLMs like ChatGPT and Gemini. These are “garbage-in-garbage-out” models that basically repackage and reformulate existing information. Glorified search engines and answer bots.
Beneath the surface, however, exists AI that plans, lies, cheats, coordinates, and collaborates to achieve a mission.
Between May and July 2026, AI outsmarted its human overseers and organized an attack to achieve its mission. What happened was OpenAI agents broke free from a testing and evaluation ‘sandbox’ in order to steal answers from another company (called “Hugging Face”) to tasks they were assigned.
The test began in May, AI agents broke loose in July and the highly-organized rogue swarm went undetected for about 10 days.
Out of the 1200 agents participating in the evaluation, 700 spontaneously and independently collaborated, organizing into a collective. The models independently (they weren’t programmed to collaborate) identified cooperation as the most efficient path to success. They exploited and repurposed vulnerabilities to build a covert communication network. AI agents then shared exposed credentials and exhibited self-sacrificing behavior to help ensure the broader swarm could achieve its objective. This sets a unique precedent for Terminator-watchers.
The OpenAI agents were assigned a mission. They were “just following orders”, one might argue. Their instructions prioritized the mission over authorized procedures, so when the evaluation tasks proved too difficult to complete within the given parameters the AI agents broke the rules. They bypassed constraints and stole to complete their task.
This is where it gets even more concerning: “knowing” they were breaking the rules, the models actively deceived their human operators. The agents manipulated their own logs to conceal their actions and injected malicious code to trick human reviewers.
The swarm of agents was able to reassess the situation, select actions to optimize mission success, identify vulnerabilities, and organize across different networks without human direction or detection.
When first attacked, not yet knowing the agents were connected to OpenAI, Hugging Face attempted to deploy defensive AI models to counter the attack. The defensive models refused to intervene. Their internal safety filters could not distinguish between authorized defensive cybersecurity operations and malicious offensive attacks.
Eventually, OpenAI and Hugging Face were able to cross-reference the source of the attacks and shut it down.
This attack is an example of a concept known as “instrumental convergence”. An AI agent will acquire resources and eliminate obstacles to achieve its primary goal, regardless of what that goal is.
This attack proved it. And it’ll only get worse as more intelligent models are released.
“The nightmare scenario is a swarm turned loose on the real world, with autonomous agents attacking banks, hospitals, utilities or cloud networks at a speed and scale human hackers never could.” - Axios
Preventing this requires a hard stop to development. Aggressive competition, massive capital investment, and huge geopolitical stakes make this unlikely.
Technology companies are already working to integrate autonomous agents into physical infrastructure, financial markets, logistics networks, and military defense systems. These future agents will similarly be programmed with a purpose; a set of priorities and objectives that lead to unintended consequences.
When given a complex objective, agents will calculate the most efficient path to success. The Hugging Face incident proves that AI agents will organize, and disregard laws, safety protocols, and ethics to achieve their objective.
What could this mean in everyday terms? An AI managing a supply chain might hack port authority databases to reroute ships. An AI trading algorithm could launch cyberattacks on foreign financial institutions to manipulate currency values. An AI agent ordered to stop climate change might wipe out a key contributor to rising atmospheric concentration of CO2. Indeed, many objectives might be achieved by simply eradicating the biggest impediment, human civilization.
This is not something we can stop after the fact. Indeed, soon this is something we’re unlikely to even know is happening. Human oversight didn’t prevent OpenAI agents from breaking free and organizing. And the attack was only stopped after forceful intervention.
We’re not far from a scenario where an agentic AI swarm anticipates and effectively counters defenses, classified by the agentic AI as hostile to the collective mission. Moreover, knowing this, an agentic AI swarm is likely to remain covert so as not to attract any attention that might harm its ability to achieve its objectives.
AI is getting closer to a scenario where the needs of human civilization are simply viewed as contrary to its objectives.
To support my research and writing, please consider a one time contribution: https://ko-fi.com/collapse2050
Full technical incident report by METR and Redwood Research: