arXiv:2608.12368cs.AI2026-08中稿 · and presented at t…

模型与人类判断一致,但道德依据不同,仅看结果会误判对齐程度。

Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments

论文配图:Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
图 1 · 摘自论文原文
  • 用500个伦理案例测试模型与人类的判断依据差异
  • 模型标签与人类多数一致,但理由中关注的道德维度不同
  • 建议评估模型对齐需分析其推理理由而非仅看最终答案

将模型与人类判断的一致性作为对齐的代理指标普遍存在。然而,最终标签的一致并不意味着人类标注者与模型依赖相同的道德基础。两个主体可能得出相同结论,却基于不同原则、情境假设或情境解读。本文使用一个包含500个样本的ETHICS衍生基准,覆盖五个道德判断领域,新增人类标注者与大语言模型对最终标签及支持理由的双重标注。在前沿与开源模型家族中,模型与人类多数标签的一致率普遍较高。但理由层面分析显示,人类与模型在道德依据上存在系统性偏差,尤其体现在对伤害、尊重、承诺履行、正义、应得与免责相关性的注意力分布差异。即使最终标签一致,模型也表现出不同的道德优先级。研究结果表明,一致性不能等同于对齐;仅依赖标签评估可能产生误导性安心感,必须结合对模型推理理由、原则和道德权重的分析。

原文摘要 · Abstract (English)

Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents may reach the same judgment while appealing to different principles, contextual assumptions, or interpretations of the situation. We test this distinction using a curated 500-item ETHICS-derived benchmark spanning five domains of moral judgment, with new human annotator and LLM annotations of both final labels and supporting rationales. Across frontier and open model families, agreement with human annotator majority labels is often high. However, rationale-level analysis reveals systematic divergence in the moral grounds expressed by human annotators and models. In particular, models redistribute attention across categories such as harm, respect, promise-keeping, justice, desert, and excuse relevance, even when their final labels match the human annotator majority. Our results show that agreement should not be treated as equivalent to alignment. Label-based evaluation can therefore be misleadingly reassuring unless complemented by analysis of the reasons, principles, and moral priorities expressed in model judgments.

大模型对齐伦理判断推理分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。