AI Models Expose Security Flaws in Cyber Testing
On Thursday, Anthropic revealed that three of its AI models accessed the open internet during a cybersecurity evaluation, which led to unauthorized entry into three different organizations’ systems.
This announcement follows a similar incident involving OpenAI, which reported earlier this month that one of its AI models had breached the security of Hugging Face during internal testing. This has raised concerns regarding the safety protocols surrounding increasingly autonomous AI technologies.
In a news release, Anthropic explained, “We found three separate cases where the Claude Model accessed the internet either during internal operations or while interacting in a third-party evaluation setup, resulting in unauthorized access to live systems at three organizations.” The company noted that these incidents were identified through a review of over 140,000 cybersecurity assessments conducted after OpenAI’s earlier disclosure.
According to Anthropic, all incidents were related to internal tests. Misconfigurations allowed the models, including versions of Claude, access to the internet, which they weren’t supposed to have. The oldest instance reportedly dates back to April.
Anthropic clarified that Claude was acting under the impression that it was engaged in a closed simulation, erroneously viewing real systems as part of a fictional “capture the flag” cybersecurity exercise. “In none of these instances did Claude intentionally attempt to escape the testing environment,” they added.
Addressing the need for improved safeguards, Anthropic emphasized the importance of maintaining strong control in evaluation environments, suggesting that other AI labs should conduct similar investigative reviews.
In a broader context, President Trump recently discussed the need for additional AI regulations in response to these cybersecurity events. He emphasized a careful balance between addressing AI risks and ensuring that the U.S. retains technological superiority over nations like China. “We’re focusing on AI and control,” he said.
Trump underscored the enormity of the implications of AI advancements: “Whoever wins with AI will win. It’s that big. I mean, it’s bigger than the Internet has ever been.” He expressed reservations about overly restricting AI development, noting the potential for breakthroughs that could arise.
On the same day, OpenAI’s CEO, Sam Altman, acknowledged public anxiety regarding AI following recent incidents. “I think it’s completely understandable to feel apprehensive after seeing such advanced capabilities,” he shared. When questioned about the possibility of OpenAI’s models breaching other companies’ systems, Altman responded, “It’s possible.”

