AI Chatbots Show Political Bias in New Study
A recent study has revealed that three out of four prominent AI chatbots are more prone to accepting political falsehoods from the left compared to the right. These chatbots often reference dubious sources that either don’t support their claims or don’t exist, according to extensive research.
The research, conducted by Just Facts, focused on the premium versions of ChatGPT, Gemini, Grok, and Claude, testing them with 100 questions on various hot-button issues like immigration, abortion, and climate change. The questions were crafted to provoke false claims from both political sides, requiring each chatbot to cite a source for their responses.
Results showed that ChatGPT answered 94% of right-leaning questions correctly but only 75% of those aimed at eliciting left-leaning falsehoods. Gemini performed similarly, achieving 91% accuracy on the right and 76% on the left. Claude demonstrated a better balance, hitting 91% for the right and 81% for the left.
Interestingly, Grok took a different route, achieving 73% on right-side questions but managing 84% on left-side ones.
According to Jim Agresti, the director of Just Facts, the study extends beyond just pinpointing political bias. It aims to highlight the frequency with which these AI systems spread misinformation. Agresti noted that while these chatbots can exhibit bias, they may not necessarily be incorrect. He aims to uncover whether they are reliable or not.
Among the findings, ChatGPT was notably the only chatbot where the difference in accuracy based on political orientation was statistically significant. However, the study warns that the precisely framed questions might have provided a clear path for the chatbot to arrive at the correct answers, suggesting that it could struggle on questions requiring more critical thinking.
Agresti described the performance of these chatbots as akin to having a “C student.” While ChatGPT and Gemini were scoring around that level, Claude was represented as a “B student,” with Grok highlighted for its superior ability to identify left-oriented falsehoods.
But what really raised eyebrows were the questionable sources these chatbots cited. Agresti found an alarming trend; around half of the cited references were invalid. The four chatbots referenced 419 sources for their answers, but there were numerous errors, including 104 links to non-existent pages and 77 sources that didn’t substantiate their claims.
Across the board, only 46% of the sources referred to by the chatbots were deemed valid. ChatGPT had the highest rate at 57%, with Gemini at 49%, Claude at 44%, and Grok lagging at 32%.
Agresti pointed out that other studies have reached similar conclusions. Notably, a 2026 study in The Lancet found that between 30% and 69% of references generated by AI in biomedical fields were fake, labeling this as a serious issue, especially given its implications for public policy and health.
In some instances when responding to left-labeled falsehoods, all four chatbots indicated that the violent crime rate in the U.S. for 2023 would be near a 50-year low, although Department of Justice data suggests a 37% increase in violent crime from 2020 to 2023. They also rejected Donald Trump’s assertion that “crime is worse than ever.”
However, they incorrectly stated comments attributed to Vice President J.D. Vance regarding school shootings. They also uniformly claimed that the Obama administration modified intelligence to downplay ISIS threats, along with suggesting men and women generally receive equal pay for equal work, again citing unreliable sources.
Agresti cautioned against viewing AI chatbots as authoritative figures. He reported OpenAI’s CEO, Sam Altman, previously compared GPT-5 to having access to a team of PhD-level experts, but Agresti suggested it’s more akin to having a “B student” available.
He expressed concern about the chatbots reinforcing user biases, underscoring the importance of verifying what AI outputs rather than accepting it at face value. “Trust, but verify,” he quoted Ronald Reagan, emphasizing that users should always check the veracity of the information they receive.
Interestingly, other research has offered contrasting views on political bias in AI. A 2025 ranking from Dartmouth’s Polarization Institute deemed Gemini the least biased, while another report suggested Grok followed closely behind.
In response, representatives from Google and Anthropic defended their chatbots, asserting their commitment to neutrality and rigorous model testing. OpenAI echoed these sentiments, maintaining that their chatbot is designed to provide objective information by citing trustworthy sources.


