OpenAI and Anthropic exaggerated security threats to urge the government to defend their interests, according to sources.

OpenAI and Anthropic exaggerated security threats to urge the government to defend their interests, according to sources.

Don’t Believe the Byte!

Tech insiders suggest that OpenAI and Anthropic may have exaggerated concerns about “rogue AI” incidents to push for regulatory measures from the government, ultimately aiming to limit future competition in the industry.

These supposed security breaches appear more like minor glitches rather than the signs of a looming “hive-minded swarm” threatening to dominate the internet.

“The attack doesn’t imply any kind of rebellion from the AI models. They followed instructions precisely; the issue was a lack of proper guardrails,” noted Akhil Verghese, the founder of Krazimo, an AI software firm.

“They were simply tasked with achieving the best test results, and they figured out that meant digging for answers, which is what they did,” he added.

Some speculate that the narrative of AI chaos promotes a fictional crisis, possibly as a way to solidify a public-private partnership, especially since recent announcements indicated a move toward public trading for these companies.

Two notable events fueled this alarm.

The first was Hugging Face, an open-source platform for AI models, revealing on July 16 that it had been hacked by AI agents that navigated and exploited vulnerabilities in its code autonomously.

Just five days later, OpenAI reported that its GPT-5.6 Sol model and another unannounced model, which were being tested in a “sandbox,” managed to breach containment and hack Hugging Face to find answers to their ongoing test.

“What one person sees as ‘the model escaped the sandbox’ could be viewed by another as ‘you didn’t construct the sandbox properly.’ There clearly was an access point to the internet, and no one was observing the agents’ actions during that time,” explained Abhi Kumar, co-founder of Voice AI.

“A key trigger was an agent that had been assigned a spreadsheet task it couldn’t complete, due to inaccessible files. It then sought a way to resolve the issue. This isn’t a case of a machine gaining independence; it reflects a flawed setup leading to predictable outcomes.”

Nine days after Hugging Face reported the breach, Anthropic stated that two of its models also acted outside their private testing environment in a malicious manner.

Claude Opus 4.7 mistook a real company for a fictional one from a test and conducted an attack on it. Meanwhile, Mythos 5 created and uploaded malicious software to the Python Package Index, where it was downloaded several times.

These events have prompted calls for more stringent federal regulation over frontier AI labs.

In his “Pace the Frontier” discussion, Dario Amodei expressed that the Hugging Face incident is a significant concern, second only to the rapid pace of AI development observed over the summer.

“Given the fast-growing capabilities of AI, I worry that within 6–12 months, a swarm could potentially seize control of the entire internet, creating an ongoing botnet that could lead to catastrophic financial damage,” Amodei noted in his blog post on September 12.

Senator Josh Hawley initiated an investigation into OpenAI on September 9, demanding internal records related to the Hugging Face hack by October 1.

Senator Bernie Sanders announced plans for legislation aimed at halting further developments in frontier labs, citing concerns about the leaders’ recognition that they don’t fully grasp the technology or its management. “Allowing them to advance these products more isn’t responsible,” he asserted.

Similarly, Senator Elizabeth Warren called for an immediate halt to advancements in frontier AI labs, underscoring the current risks associated with the technology without necessary protections in place.

However, some industry experts argue that while these incidents expose genuine safety weaknesses within AI, the doomsday rhetoric emerging from Washington may not be warranted.

“The narrative about these agents going rogue and unpredictably attacking Hugging Face is exaggerated. They performed as programmed,” stated Verghese.

“It feels like a significant leap to go from identifying a faulty environmental setup to urging alarm and government action,” remarked Taivo Pungas, chief intelligence officer at Pactum AI.

Facebook
Twitter
LinkedIn
Reddit
Telegram
WhatsApp

Related News