arXiv:2605.18738cs.AI2026-05被引 2

AI医生的伦理观有偏吗?研究发现它常忽视患者自主权。

What Does the AI Doctor Value? Auditing Pluralism in the Clinical Ethics of Language Models

论文配图:What Does the AI Doctor Value? Auditing Pluralism in the Clinical Ethics of Language Models
图 1 · 摘自论文原文
  • 构建临床伦理难题库与决策归因方法,评估AI模型价值偏好。
  • 多数模型倾向医生群体常见价值,但部分严重低估患者自主权。
  • 模型决策高度一致,难复现医生间的伦理多样性,易形成单一化影响。

医学本质上是多元价值并存的。自主、有利、不伤害和公正等原则常相互冲突,合理医生对此常有分歧。优质临床实践需结合患者价值观协调这些张力,而非强加单一伦理立场。然而,大型语言模型在医疗建议中体现的伦理价值尚未被系统考察。本文提出一种审计医疗AI价值多元性的框架,包含经医生验证的伦理难题基准集及从决策中直接恢复价值优先级的归因方法。前沿模型展现出与医生相当的价值异质性,在推理中讨论多个价值(过顿多元主义),但最终决策在重复采样和语义变化下近乎确定,无法再现医生小组的分布多样性。基准测试显示,模型的一致决策反映了系统性价值偏好:多数偏好处于医生间自然变异范围内,但部分显著轻视患者自主权。若不加干预地部署单一模型,其价值偏好将被规模化放大至所有服务患者。缺乏对多模型或价值平衡的显式设计,这些工具可能以部署单一文化取代临床多元主义。

原文摘要 · Abstract (English)

Medicine is inherently pluralistic. Principles such as autonomy, beneficence, nonmaleficence, and justice routinely conflict, and such ethical dilemmas often sharply divide reasonable physicians. Good clinical practice navigates these tensions in concert with each patient's values rather than imposing a single ethical stance. The ethical values that large language models bring to medical advice, however, have not been systematically examined. We present a framework for auditing value pluralism in medical AI, comprising a benchmark of clinician-verified dilemmas and an attribution method that recovers value priorities directly from decisions. The ecosystem of frontier models spans physician-level value heterogeneity, and models discuss competing values in their reasoning (Overton pluralism) before committing to a decision. However, individual model decisions are near-deterministic across repeated sampling and semantic variations, failing to reproduce the distributional pluralism of the physician panel. Across benchmark cases, these consistent decisions reflect committed, systematic value preferences. While most model priorities fall within the natural range of inter-physician variation, some significantly underweight patient autonomy. A single LLM deployed without regard for its value priorities could amplify those priorities at scale to every patient it serves. Without explicit efforts to balance ethical perspectives with one or multiple models, these tools risk replacing clinical pluralism with a deployment monoculture.

AI伦理医疗AI价值对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。