arXiv:2604.21871cs.CL2026-04ACL被引 1

LLM做道德判断时,选对但不接地气,像按规则办事的机器。

Machine Behavior in Relational Moral Dilemmas: Moral Rightness, Predicted Human Behavior, and Model Decisions

论文配图:Machine Behavior in Relational Moral Dilemmas: Moral Rightness, Predicted Human Behavior, and Model Decisions
图 1 · 摘自论文原文
  • 用举报人困境测试模型在关系亲疏与罪行严重性下的决策
  • 模型选对但预测人类会更讲情面,说明它懂人情却不用
  • 适合研究大模型伦理对齐、社会智能的团队参考

人类道德判断受人际关系影响,具有情境依赖性。随着大语言模型(LLMs)越来越多作为决策辅助系统使用,理解其是否具备这些社会细微差别至关重要。我们通过改变犯罪严重性和关系亲密度两个维度,在举报人困境中评估模型行为。研究涵盖三个视角:(1) 道德正确性(规范性准则),(2) 人类行为预测(描述性社会期待),(3) 模型自主决策。分析推理过程发现三者存在明显分歧:尽管道德正确性始终以公平为导向,但对人类行为的预测随关系亲密度上升显著转向忠诚。关键的是,模型决策与道德正确性一致,而非其自身对人类行为的预测。这表明模型决策更依赖僵化的规范性规则,而非其内部世界模型所体现的社会敏感性,暴露出真实部署中可能产生重大偏差的差距。

原文摘要 · Abstract (English)

Human moral judgment is context-dependent and modulated by interpersonal relationships. As large language models (LLMs) increasingly function as decision-support systems, determining whether they encode these social nuances is critical. We characterize machine behavior using the Whistleblower's Dilemma by varying two experimental dimensions: crime severity and relational closeness. Our study evaluates three distinct perspectives: (1) moral rightness (prescriptive norms), (2) predicted human behavior (descriptive social expectations), and (3) autonomous model decision-making. By analyzing the reasoning processes, we identify a clear cross-perspective divergence: while moral rightness remains consistently fairness-oriented, predicted human behavior shifts significantly toward loyalty as relational closeness increases. Crucially, model decisions align with moral rightness judgments rather than their own behavioral predictions. This inconsistency suggests that LLM decision-making prioritizes rigid, prescriptive rules over the social sensitivity present in their internal world-modeling, which poses a gap that may lead to significant misalignments in real-world deployments.

大模型伦理道德判断社会智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。