Anthropic Discloses How It Prevented ‘Kamikaze Drone’ Swarms, Biological Weapons, and More

Anthropic Discloses How It Prevented ‘Kamikaze Drone’ Swarms, Biological Weapons, and More

Anthropic’s Efforts to Combat AI Misuse

Anthropic has released a new report highlighting its initiatives to prevent the misuse of its Claude artificial intelligence models by malicious entities.

The report reveals various cases, including fraudulent dating apps and surveillance systems aimed at monitoring dissenters, featuring insights from Anthropic’s Threat Intelligence report. It categorizes the bad actors as “suspected state-sponsored groups, financially motivated criminals, commercial spyware vendors, state propaganda institutions, and politically motivated individuals.”

According to the findings, misuse of Claude models — specifically Haiku, Sonnet, and Opus — was prevalent, while their more advanced models, Fable and Mythos, were mostly untouched, except for one incident involving an attempt to steal AI model data through a method called “distillation.”

We’re publishing our most detailed threat intelligence report to date. It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them. We disrupted every operation in the report,…

This report comes on the heels of a recent resignation from an Anthropic researcher who expressed concerns about the potential dangers of AI, suggesting it could have catastrophic consequences within the decade.

I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.

The company noted that it had banned a number of Chinese accounts linked to state security forces that used Anthropic’s AI for what they termed “stability maintenance,” a phrase used by the Chinese Communist Party to describe suppressing dissent. Anthropic also identified an actor believed to be connected to a municipal cyber police unit that utilized Claude to operate a domestic surveillance system, pinpointing ten Chinese citizens as targets.

Among those being surveilled were pro-democracy advocates in Hong Kong, those passionate about commemorating the Tiananmen Square massacre, as well as Uyghur supporters and Western human rights organizations.

Following the UK’s handover of Hong Kong to China in 1997, the Chinese government has gradually tightened its hold over the region, restricting freedom of speech and the press, as per the Council on Foreign Relations.

The Chinese Embassy in Washington, D.C., has yet to respond to requests for comment regarding these allegations.

Anthropic also reported disrupting attempts to utilize AI for developing biological weapons but could not confirm if the motivations behind this research were malicious or legitimate.

“You are not seeing someone in a comic book kind of way say, ‘Hey, I want to build a biological weapon to kill everybody,’” stated Jacob Klein, head of threat intelligence at Anthropic, in comments to The New York Times.

The report further suggested that complex attacks do not necessarily require equally sophisticated attackers.

One actor identified as GTG-20006 is thought to have links to Russian state-sponsored espionage and has targeted governmental ministries, defense and intelligence agencies, embassies, think tanks, and defense contractors. This actor seems particularly focused on Ukrainian military drone technologies and supply chains. It remains uncertain whether GTG-20006 is an individual, a group, or a collective.

Reportedly, this supposed Russian operation involved using AI to exploit hotel WiFi networks, from which malware was sent to compromise targets’ laptops and smartphones.

Anthropic also highlighted a “likely freelance Russia-based” threat actor involved in plans to create a “kaze drone swarm,” utilizing Claude Code for testing purposes. This group may have connections with the Russian Academy of Sciences rather than being directly linked to state actors.

In another noteworthy incident, a Chinese entity was found attempting to develop software capable of detecting, jamming, or potentially deceiving radar and communications systems of adversaries.

The report accused multiple Chinese laboratories—including Alibaba, Moonshot AI, DeepSeek, Z.ai, Xiami, and MiniMax—of engaging in distillation, a term Anthropic uses to describe unauthorized efforts to replicate a model’s functionality. It indicated that such practices often involve deceit, such as fraud and stolen credentials.

Moreover, it was discovered that Moonshot AI, known for its Kimi AI models, had been redirecting customer requests to Claude instead of processing them through Kimi.

After introducing Kimi K3, which generated immense demand, Moonshot AI paused new subscriptions. The Trump administration later accused the company of illicit distillation that involved American AI models to create Chinese versions.

At the time, the Chinese Embassy responded, arguing that AI development in China stemmed from a stronger foundation in science and technology and called for collaborative efforts to ensure the positive advancement of AI globally.

Anthropic expressed hopes that the report would help governments and civil society gain better insights into AI safety and evolving threats.

“Sophisticated and persistent threat actors continuously test our safeguards and try to circumvent the technical measures we use to detect and prevent misuse. We’ll continue to evolve our safeguards and coordinate with our partners to improve our ability to detect, disrupt, and prevent future misuse,” the report concluded.

“We hope that the findings in this report will help other developers recognize similar patterns on their own platforms, give governments and civil society a clearer view of how emerging threats take shape, and strengthen collective defenses,” it added.

Facebook
Twitter
LinkedIn
Reddit
Telegram
WhatsApp

Related News