研究大模型在不同情境下的道德判断变化,发现其敏感性与人类不一致。
Between Rules and Reality: On the Context Sensitivity of LLM Moral Judgment
- 构建含情境变量的道德困境数据集,模拟人类心理影响因素
- 22个模型均表现出情境敏感性,多数倾向违反规则
- 提出激活控制方法,可调节模型的情境敏感程度
人类的道德决策高度依赖情境。然而,当前大模型道德研究多基于固定场景。本文引入Contextual MoralChoice数据集,包含道德心理学中已知会改变人类判断的三类情境变量:功利主义、情感和关系因素。评估22个大模型发现,几乎所有模型都表现出情境敏感性,其判断趋向于违反规则的行为。与人类问卷调查对比显示,模型与人类对不同情境变量的反应差异显著;即使模型在基础情况下与人类判断一致,其情境敏感性仍可能不匹配。这引发对控制情境敏感性的思考,本文提出一种激活操控方法,可稳定提升或降低模型的情境敏感性。
原文摘要 · Abstract (English)
A human's moral decision depends heavily on the context. Yet research on LLM morality has largely studied fixed scenarios. We address this gap by introducing Contextual MoralChoice, a dataset of moral dilemmas with systematic contextual variations known from moral psychology to shift human judgment: consequentialist, emotional, and relational. Evaluating 22 LLMs, we find that nearly all models are context-sensitive, shifting their judgments toward rule-violating behavior. Comparing with a human survey, we find that models and humans are most triggered by different contextual variations, and that a model aligned with human judgments in the base case is not necessarily aligned in its contextual sensitivity. This raises the question of controlling contextual sensitivity, which we address with an activation steering approach that can reliably increase or decrease a model's contextual sensitivity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。