arXiv:2609.05437cs.AIcs.CY2026-09

测试大模型对社会规则背后隐含反应的推理能力,发现其过于严苛。

Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models

论文配图:Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models
图 1 · 摘自论文原文
  • 构建双维度评估框架,分析模型对违规者自控与旁观者干预的判断
  • 6个模型均高估惩罚,人类在远距离关系中更倾向沉默,模型却反向预测
  • 适合关注AI社会认知偏差、伦理对齐的研究者和政策模拟应用者

以往人工智能对齐研究主要聚焦第一层社会规范(如‘不要偷窃’),但社会智能还依赖于对规范执行者及方式的预判(如公开羞辱或监禁)。这类二阶预期称为元规范,决定人们如何应对规则被破坏。本文提出新评估框架,从情感评价与行为反应两维衡量大语言模型的元规范推理能力,并设计预测违规者自我约束与旁观者干预的新分类任务。我们发布多视角数据集NormReact,包含450个规范违背场景,人工标注了违规者性别与观察者亲疏关系下的情绪与行为反应。结果显示:六种当前主流模型普遍高估负面制裁,当人类预期不作为时,模型仍预测惩罚;且随着社会距离增加,模型与人类判断的一致性下降。这表明,在冲突调解、政策模拟等对齐敏感领域,当前AI可能生成扭曲的社会调控图景——过度强调惩罚,低估实际中的宽容、克制与关系调节。

原文摘要 · Abstract (English)

Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., `do not steal'). However, social intelligence depends not only on norm recognition, but also on anticipating who will enforce it and how (e.g., public shame or even imprisonment). These second-order expectations, known as metanorms, govern how people respond when social rules are broken. We introduce a novel framework for evaluating metanorm reasoning in Large Language Models (LLMs) along two dimensions: emotional appraisal and behavioral response, and propose new classification tasks, namely, predicting self-regulation in violators, and other-regulation in observers. We release a multi-perspective dataset, NormReact, of 450 norm violation scenarios, hand-annotated for emotions and behavioral responses across norm violators' gender and observers' social closeness. Current LLMs portray a harsher social world: across six models, they overpredict negative sanctions where humans would expect inaction, and alignment with human judgments deteriorates as social distance increases. These findings suggest that AI systems in norm-sensitive domains from conflict mediation to policy simulation, may risk producing a distorted picture of social regulation: one that over-represents punishment and under-represents the tolerance, restraint, and relational calibration that characterize actual norm enforcement in real world.

社会推理元规范大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。