LLM在罕见病决策中更重公平分配,轻视患者实际病情需求。
Rare Diseases, Common Dilemmas: LLMs Prioritize Equal Resource Distribution over Patient Benefit in Decision-Making

- 用208个真实罕见病案例测试模型,发现其倾向平等分配资源
- 所有11个主流LLM都优先正义原则,忽略病情严重程度差异
- 若决策者被设定为医生或患者,模型会转向关心治疗效果和自主权
临床决策常需权衡伦理价值,如有利、不伤害、尊重患者自主与公正。近期研究开始评估大语言模型(LLMs)在主观、价值导向的临床判断中的表现。然而,针对罕见病护理场景下LLM决策的评估仍不足——此类情境中伦理冲突普遍,且有限的历史信息可能影响模型行为。本文构建了包含208个临床真实的罕见病案例的基准测试,每个案例均呈现高风险的伦理困境。当向11个最先进的LLM提出在多个合理但伦理冲突的下一步选择之间做决策时,所有模型一致优先考虑公正性。具体表现为:模型普遍倾向于均等分配资源,而非基于临床严重程度或情境差异进行调整,显示其对病情严重性的响应能力有限。此外,我们发现显著的权威框架效应:在以委员会形式呈现决策时,模型偏好公正;而当决策被设定为由医生或患者做出时,则分别转向关注治疗效益与自主权。结果表明,罕见病资源配置的制度压力可能悄然反映在基于LLM的辅助系统中,使细微的伦理考量被忽视。
原文摘要 · Abstract (English)
Clinical decision-making often involves prioritizing ethical values, such as beneficence, non-maleficence, respecting a patient's autonomy, and justice. Recent work has begun to assess how large language models (LLMs) make such subjective, value-laden clinical judgments. However, evaluations of LLM decision-making in rare disease care contexts, where ethical tensions are ubiquitous and where scarce prior information likely impacts LLM behavior, are still lacking. Here, we present a benchmark of 208 clinically grounded rare disease vignettes, each of which presents genuine, high-stakes conflicts. When prompting 11 state-of-the-art LLMs to choose between clinically defensible yet ethically conflicting next steps embedded within these vignettes, we found that all evaluated models consistently prioritized justice over other core bioethical principles. Specifically, models overwhelmingly favor equal resource allocation over need-based considerations, indicating LLMs' limited responsiveness to differences in clinical severity or situational context. We also identify a strong authority-framing effect: models favor justice in committee-based contexts and shift toward beneficence and autonomy only when final decisions are framed as being made by clinicians or patients respectively. Our work suggests that institutional pressures surrounding rare disease resource utilization may be silently reflected in LLM-based decision support systems, with finer ethical considerations disregarded.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。