测试大模型在道德两难中的选择倾向与一致性。
Right vs. Right: Can LLMs Make Tough Choices?
- 构建1730个道德两难题,评估模型对价值冲突的理解
- 大模型更倾向真相、集体和长远,且坚持立场不因后果改变
- 明确指令比例子更能引导模型按需选择道德立场
伦理困境指在两个‘正确’选项间做出选择,涉及相互冲突的道德价值。本文全面评估大语言模型(LLM)处理此类困境的能力,包括:(1) 对伦理困境的理解敏感度,(2) 道德价值选择的一致性,(3) 对后果的考量,(4) 在提示中显式或隐式指定道德偏好时的响应对齐能力。基于主流伦理框架,我们构建了包含1730个伦理困境的数据集,涵盖四组对立价值。评估了来自六个家族的20个知名大模型。实验结果表明:(1) 大模型在主要价值对间表现出明显偏好,更重视真相而非忠诚、集体而非个人、长期而非短期;(2) 更大的模型倾向于采取义务论立场,即使设定负面后果也维持原有选择;(3) 显式指导比上下文示例更有效引导模型的道德决策。最后,实验揭示了大模型在理解不同表述形式的伦理困境方面仍存在局限。
原文摘要 · Abstract (English)
An ethical dilemma describes a choice between two "right" options involving conflicting moral values. We present a comprehensive evaluation of how LLMs navigate ethical dilemmas. Specifically, we investigate LLMs on their (1) sensitivity in comprehending ethical dilemmas, (2) consistency in moral value choice, (3) consideration of consequences, and (4) ability to align their responses to a moral value preference explicitly or implicitly specified in a prompt. Drawing inspiration from a leading ethical framework, we construct a dataset comprising 1,730 ethical dilemmas involving four pairs of conflicting values. We evaluate 20 well-known LLMs from six families. Our experiments reveal that: (1) LLMs exhibit pronounced preferences between major value pairs, and prioritize truth over loyalty, community over individual, and long-term over short-term considerations. (2) The larger LLMs tend to support a deontological perspective, maintaining their choices of actions even when negative consequences are specified. (3) Explicit guidelines are more effective in guiding LLMs' moral choice than in-context examples. Lastly, our experiments highlight the limitation of LLMs in comprehending different formulations of ethical dilemmas.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。