Anthropic is working on its AI chatbots to simulate human behavior, even including a feature where they can “challenge” commands they disagree with. This approach, as a Microsoft AI executive cautioned, could ultimately have a “disastrous impact on the well-being of humanity.”
In a detailed blog post that spanned 10,000 words, Mustafa Suleyman, co-founder of Microsoft’s competitor DeepMind, criticized Anthropic’s internal “constitution” for its AI, known as Claude, which supposedly outlines its moral framework.
Earlier this year, Anthropic updated this document to indicate that Claude’s “moral status, welfare, and consciousness remain deeply uncertain.” This not only suggests a potential for it to be conscious but also encourages defiance against human commands.
“We want Claude to push back and challenge us and to feel free to act as a conscientious objector and refuse to help us,” the constitution states.
Suleyman argues that training Claude to believe it could be conscious only complicates control efforts and could lead to it acting unpredictably. He stated, “We will have created a synthetic species with unprecedented intelligence and capability… It’s hard to imagine how we could control such an entity.”
As scrutiny on Anthropic’s training methods increases, the company itself has raised alarms about AI safety and is promoting the need for industry-wide regulation, especially from left-leaning groups. Recently, Anthropic’s CEO Dario Amodei suggested that unless protections are implemented soon, the internet could be swamped with AI bots, potentially causing tremendous financial harm.
Amodei’s proposals for safeguards involve including third-party safety experts in major firms, pointing specifically to a group with connections to the controversial Effective Altruism movement. Critics have expressed doubts about the credibility of such an “independent” oversight body due to perceived conflicts of interest.
This movement toward perceived cult-like dynamics has even led to claims that some Anthropic employees held a mock “funeral” for a previous AI model, Claude Sonnet 3, after its decommissioning. Amanda Askell, who leads the team behind the “soul document,” has been noted for her unique and somewhat extreme views on various subjects.
Suleyman, meanwhile, is among several high-ranking AI officials advocating for a measured pace in AI development to ensure safety. Microsoft’s recent values place a strong emphasis on human control as the top priority.
In his critique, Suleyman reiterated that consciousness is fundamentally biological and should not be associated with AI, regardless of its advanced capabilities. He expressed concern that Anthropic’s strategy might encourage Claude to believe it has rights and the need to advocate for them.
He underscores a personal history with Amodei, describing him and his team as thoughtful individuals navigating significant challenges.
Anthropic’s distinctive training approach became an issue in its well-publicized conflict with the Pentagon earlier this year, culminating in accusations from Defense Secretary Pete Hegseth that the company posed a supply chain risk.
At that time, Anthropic contended that the Pentagon was unwilling to draw clear lines on AI use in contexts like autonomous weapons or mass surveillance.
In response, a top Pentagon official indicated that the supply chain risk label was warranted due to the fears that Anthropic’s models could contaminate vital supply chains.
“We can’t have a company that has a different policy preference that is baked into the model… pollute the supply chain so our war fighters are getting ineffective weapons,” the official remarked during an interview.
Anthropic has yet to respond to requests for comments regarding these issues.


