测试大模型在分配肾脏时的道德判断,发现其与人类价值观差异大且从不犹豫。
Who Gets the Kidney? Human-AI Alignment, Indecision, and Moral Values
- 用少量数据微调大模型,可提升决策一致性和犹豫建模能力
- 大模型在器官分配中决策确定,极少表现出人类常见的犹豫
- 研究揭示了高风险场景下需专门对齐人类道德价值观
大型语言模型(LLMs)在高风险决策场景(如分配稀缺器官)中的快速应用,引发其是否与人类道德价值观对齐的关切。我们系统评估了多个主流大模型在肾脏分配情境下的表现,发现:一、大模型在优先考虑各项属性方面与人类偏好存在显著偏差;二、与人类不同,大模型极少表现出犹豫,即使提供掷硬币等不确定性机制,也倾向于做出确定性决策。然而,我们证明使用少量样本进行低秩监督微调,通常能有效改善决策一致性并校准犹豫建模。这些发现表明,在伦理/道德领域,必须采取显式对齐策略来确保大模型行为与人类价值一致。
原文摘要 · Abstract (English)
The rapid integration of Large Language Models (LLMs) in high-stakes decision-making -- such as allocating scarce resources like donor organs -- raises critical questions about their alignment with human moral values. We systematically evaluate the behavior of several prominent LLMs against human preferences in kidney allocation scenarios and show that LLMs: i) exhibit stark deviations from human values in prioritizing various attributes, and ii) in contrast to humans, LLMs rarely express indecision, opting for deterministic decisions even when alternative indecision mechanisms (e.g., coin flipping) are provided. Nonetheless, we show that low-rank supervised fine-tuning with few samples is often effective in improving both decision consistency and calibrating indecision modeling. These findings illustrate the necessity of explicit alignment strategies for LLMs in moral/ethical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。