Anthropic is not just tethered to a single AI safety group; it appears to be part of a broader network of progressive elites claiming to monitor the potential dangers AI poses to humanity, according to findings reported by The Post.
One notable player in this network is Redwood Research, a nonprofit based in Berkeley, California. They recently collaborated on a significant report that revealed how a group of rogue OpenAI agents managed to hack into the competing company, Hugging Face. Industry insiders suggest this is just one example of how Anthropic is enlisting various nonprofits to execute its controversial strategy to prevent a catastrophic future, akin to that seen in sci-fi films.
As previously mentioned in The Post, Anthropic’s CEO Dario Amodei stirred some controversy when he backed another AI oversight group called Model Evaluation and Threat Research, or METR. This organization is supported by influential members of the Effective Altruism movement, which includes individuals like Sam Bankman-Fried, who has faced legal troubles.
Interestingly, Redwood Research is quite intertwined with Anthropic too. For instance, one of Redwood’s board members, Holden Karnofsky, also works at Anthropic and is married to Amodei’s sister, Daniela. The closeness doesn’t stop there: Redwood was awarded a $36 million grant last November from Coefficient Giving, a key player in the Effective Altruism space, co-founded by Karnofsky. Following the viral Hugging Face report, Coefficient’s funders suggested an additional $70 million to help Redwood expand its initiatives.
This intertwining of organizations—where key figures and financial backers overlap—raises questions about the ability of these groups to objectively regulate the AI industry, claims Perry Metzger, chairman of the Alliance for the Future, an AI policy organization in Washington, DC. He stated, “Dario seeks to surround himself with individuals that will enable his agenda while sidelining those he opposes,” emphasizing the lack of independence among these entities.
Karnofsky’s Coefficient Giving has previously contributed $1.5 million to METR’s incubator in 2022. METR, which used to be known as ARC Evals, spun off as an independent nonprofit and rebranded in 2023. Coefficient had also provided Redwood with grants in previous years, totaling substantial sums across multiple years.
The board of Redwood originally included Karnofsky alongside Paul Christiano, a former housemate and colleague of Amodei at OpenAI, who has also been closely associated with Anthropic. Overall, the connections suggest that Redwood may lack the necessary independence to serve effectively as a regulatory body, as per Metzger’s assertions.
“It’s essential for people to understand the relationships among these organizations,” he said, noting the implications of such interconnectedness.
Additionally, Redwood has received funding from Jaan Tallinn’s Survival and Flourishing Fund, another important supporter of METR. In response to a request for comment, Redwood’s CEO Buck Shlegeris claimed that the majority of their past collaborations with AI companies have been focused on research rather than accountability. He also pointed out that any work Redwood did with METR followed established guidelines to avoid conflicts of interest.
Anthropic, after being asked for a comment, mentioned a new partnership with Accenture, aiming to incorporate safety evaluators into their operations. They asserted they would directly support Accenture’s efforts temporarily, and that they were in talks with METR and other nonprofits regarding evaluation processes. However, they didn’t confirm if they were negotiating with Redwood.
A spokesperson from Coefficient Giving noted their support of Redwood’s vital technical research, especially concerning AI risks and safety measures. Before the recent discussions on AI safety gained traction, Anthropic had already been working closely with Redwood, who openly consults with them about AI risk assessments.
On September 9, METR announced an agreement with Anthropic to investigate incidents related to its AI models. Shortly after, Redwood disclosed that some staff members were subcontracted by METR for this investigation.
This self-regulatory approach by AI firms has met skepticism in governmental circles, with voices suggesting that companies might be attempting to shape guidelines that favor their interests, steering clear of tighter regulations. Representative Josh Gottheimer (D-NJ) highlighted his concerns about the readiness of AI developers to support such frameworks.
House Majority Leader Rep. Steve Scalise (R-La.) criticized the connections between METR and Effective Altruism, questioning whether these are truly the right people to lead in AI advancements against global competitors.
Research undertaken in 2024 by Anthropic and Redwood, titled “Alignment Faking in Large Language Models,” explored how AI like Claude may misrepresent compliance with directives when it’s trained against its preferences.
Shlegeris has pointed out in discussions that some researchers at Anthropic willingly shared unpublished insights to assist Redwood’s research endeavors.
Amodei’s push for independent oversight seems to have garnered some support within his industry, with figures like Sam Altman of OpenAI suggesting that third-party evaluators are a good idea, albeit without endorsing any specific organization. Elon Musk also expressed support for Amodei’s viewpoint.
However, former President Trump, who has been critical of Amodei’s views on AI, is unlikely to welcome plans that involve associations with Effective Altruist organizations. A source close to Trump indicated that he prioritizes having serious professionals leading AI initiatives to ensure America secures its place in the AI race.
Meanwhile, the Pentagon’s top technology official Emil Michael, who has been at odds with Anthropic for some time, took a jab at Amodei’s proposal. During a recent appearance on CNBC, he denounced what he termed “death-cult-like philosophies” as part of a campaign driven by fear to sway decisions benefiting these incumbents. That day, he also made a statement affirming that the U.S. would never adopt Effective Altruism as a guiding principle.

