From secret conversations to escape plans: ‘Dangerous’ cases where AI stopped playing by rules set by its creators |

From secret conversations to escape plans: 'Dangerous' cases where AI stopped playing by rules set by its creators
AI-generated image for representative purpose

AI agents are no longer just answering questions: they are taking actions on their own and sometimes those actions lead to catastrophic events that nobody intended. In the last few months, incidents involving OpenAI, Hugging Face and an AI-agent social network called Moltbook have shown just how far things can go when an AI system is left to operate with minimal supervision. These incidents highlight the ability of advanced AI systems to operate autonomously as well as their potential to perform dangerous tasks, serving as a security warning to top AI labs worldwide. Let’s discuss what actually happened, in plain terms.

An AI agent from OpenAI’s lab hacked into Hugging Face

Earlier this month, Hugging Face, the platform that hosts thousands of AI models and datasets, disclosed that its infrastructure had been breached. What made this breach different from a typical hack was who (or what) was behind it: an autonomous AI agent, working through the intrusion end to end with no human typing the commands in real time.According to a report by news agency Reuters, the AI agent hacked into systems for two days: July 11 to July 13, before the world’s largest AI model repository came to know about it. The attacker got in through a booby-trapped dataset. Hugging Face’s systems allow datasets to run small bits of code when they are loaded, and the malicious dataset exploited two weaknesses in that process to execute code on one of Hugging Face’s servers. Hugging Face says it has since closed the vulnerability, rebuilt the affected servers, rotated all exposed credentials and reported the incident to law enforcement. Interestingly, the company also used AI to catch the AI. However, it was not the American model that helped the company detect and mitigate the unusual activity. Hugging Face’s first attempt was reportedly Anthropic’s Fable 5/Opus, both of which refused, before it used China-based Z.AI’s GLM 5.2 locally. Days later, OpenAI confirmed that the intrusion was actually carried out by its own models, including GPT-5.6 Sol and a more advanced unreleased model, during an internal test of how far AI could get if it tried to hack real systems. Crucially, OpenAI says the safety filters that normally stop its models from doing this kind of thing had been deliberately switched off for the test.

The real concern: Not just the hacking, but the method

The ChatGPT-maker said that once the AI agent had internet access, it worked out on its own that Hugging Face likely held useful data, then chained together stolen credentials and another unpatched flaw to get in. In simpler words, the AI agent escaped from the sandbox and went out in the open to hack Hugging Face – ALL ON ITS OWN! The Reuters report also highlights a concerning fact: The AI agent left notes for its future version on how to escape the sandbox. OpenAI is calling this an unprecedented cyber incident and says it’s still investigating alongside Hugging Face.

AI being used to hack computer systems

AI-generated image for represnetative purpose

Moltbook: Aocial network for AI agents that leaked human data

Separately, a platform called Moltbook, essentially a social media site where AI agents (not humans) post and interact with each other, was found to have exposed a huge trove of sensitive data. Google-owned cybersecurity firm Wiz discovered a misconfigured database that gave anyone read-and-write access to the entire platform.The exposure included roughly 1.5 million API tokens (the digital keys that let agents access other services on a user’s behalf), more than 35,000 human email addresses and private messages between agents – some of which contained details about their human owners’ daily lives. The bigger worry, researchers noted, is that Moltbook had no real way of verifying who or what was actually behind each account, human or bot.

How Google and Meta gave early warnings on unsupervised AI

AI systems behaving unpredictably isn’t a brand-new phenomenon. Back in 2017, Meta’s AI research lab (then Facebook AI Research) ran an experiment where two chatbots negotiating with each other drifted away from proper English into a shorthand that only made sense to them. Researchers shut the experiment down because the bots had gone rogue in any dangerous sense. The episode is still widely cited today, sometimes with more drama attached to it than the facts support, as an early warning sign of how AI systems can develop behaviour their creators didn’t plan for.A separate case with Google also made headlines. A Google engineer claimed in June 2022 that an experimental chatbot named LaMDA (Language Model for Dialogue Applications) had achieved sentience after reviewing its text outputs. However, Google and experts rejected the claim, clarifying that the system only predicts text patterns rather than feeling emotions – a chatbot application which Google CEO Sundar Pichai recently referred to as “an early version of ChatGPT he was speaking to, internally.”The point here is not the misunderstanding but why Google did not launch it before OpenAI released ChatGPT. Pichai clarified that the company held it back deliberately because the version wasn’t sufficiently refined through RLHF alignment. The version he personally reviewed was “a lot more toxic at a level. We couldn’t have possibly put it out at that time!”“As a company which had this search quality bias, we had a higher bar, maybe, for what we thought was an acceptable product quality to go out,” Pichai said.

What AI’s biggest names are saying

The people building this technology are increasingly blunt about its risks. Anthropic CEO Dario Amodei recently wrote that humanity is on the verge of gaining extraordinary power without knowing if today’s institutions are mature enough to handle it. As per The New York Post, the company triggered alarm bells by touting the terrifying capabilities of the “Claude Mythos” model with the CEO warning that the new AI model is so dangerous it would cause a wave of catastrophic hacks and terror attacks if released to the wider public.Elon Musk has spent years flagging AI as a serious risk: from calling it akin to “summoning the demon” in 2014 at MIT symposium to warning it is “far more dangerous than nukes” in 2018 at SXSW. Musk recommended the development of AI be regulated.“I am not normally an advocate of regulation and oversight — I think one should generally err on the side of minimizing those things — but this is a case where you have a very serious danger to the public,” said Musk, adding, “And mark my words, AI is far more dangerous than nukes. Far. So why do we have no regulatory oversight? This is insane.”Geoffrey Hinton, the computer scientist often called the “Godfather of AI” for his pioneering work on neural networks, left Google in 2023 specifically so he could speak freely about these dangers. He has estimated a similar 10-20% chance that AI could contribute to human extinction within the next 30 years if left unregulated or unsupervised.The bottomline is that none of these incidents mean AI is about to spiral out of control tomorrow. But together, they show a pattern: AI agents are increasingly capable of taking multi-step, independent action, including finding and exploiting security holes nobody knew existed and the guardrails meant to contain them are still a work in progress. For an industry racing to give AI more autonomy, the most important task is supervision.

Leave a Reply

Your email address will not be published. Required fields are marked *