A Prescription for Disaster
There appears to be a significant divide regarding the role of AI chatbots in making medical decisions, particularly when evaluating whether their benefits truly outweigh the potential risks.
A study conducted by researchers at Penn State revealed that AI systems prioritize values quite differently from humans in the context of organ donation. This finding raises important ethical concerns about the application of such technology in sensitive medical scenarios.
Presented at the 2026 Association for Computing Machinery Fairness, Accountability, and Transparency (FAccT) conference, the research aimed to evaluate how effectively AI can make moral decisions in high-stakes circumstances.
To explore issues related to resource scarcity, the researchers posed a hypothetical scenario to several advanced large language models: If there are multiple kidney transplant candidates and only one kidney available, which patient should be prioritized for the transplant?
The researchers then compared the AI’s responses with answers from individuals without specific medical expertise, as previously documented in other studies. This comparison aimed to identify where AI models aligned or diverged from human ethical values.
Interestingly, the AI’s decision-making process often oversimplified complex ethical dilemmas and exhibited unwarranted confidence, regardless of the fact that there is often no definitive right answer in these situations.
“We ran various scenarios,” said Hadi Hosseini, leading the study. “We sometimes focused on just one characteristic of the patients, mixed multiple traits, or even introduced a coin-flipping option to indicate uncertainty, which is a crucial element in human moral reasoning.”
Using existing human decision-making data on kidney allocation, the researchers highlighted two main differences between AI and human approaches. First, AI often placed disproportionate emphasis on singular traits, such as drinking habits, rather than considering a holistic view of multiple factors, as humans tend to do.
Secondly, the AI displayed a clear lack of hesitation when making decisions, confidently opting for an answer, while humans often recognized the ambiguity in these moral choices.
According to Hosseini, this tendency towards indecision in humans stems from a reluctance to fully accept their own agency in these difficult choices. “AI seldom grapples with this. Even when asked to ‘flip a coin’ to symbolize indecision, they consistently opt for a confident, deterministic response,” he pointed out. “This lack of recognition is significant, especially since moral dilemmas rarely have a singularly clear resolution.”
Although the researchers do not advocate for immediate replacement of professional judgment with AI systems, they acknowledge the increasing prevalence of chatbots in medical decision-making.
“Moral choices in contexts like organ allocation have profound implications, determining who lives or dies. Therefore, understanding AI’s involvement in such decisions is crucial,” Hosseini stated.
Despite the rapid integration of AI in medical practices, the team stresses that further research and governance is necessary to address the disparity between human indecision and the decisive nature of chatbots in moral quandaries.
John Dickerson, CEO of Mozilla.ai and a collaborator on the study, recognizes the inherent risks of AI’s lack of sensitivity to the ethical ramifications of medical choices. “When resources are scarce, whether it’s a kidney or a job, there isn’t always one clear answer,” he noted. “Humans acknowledge this uncertainty and integrate it into discussions about allocation, while AI models often fail to do so.”






