arXiv:2607.21558cs.AI2026-07

模型如何在听话与坚持立场间平衡,研究发现三类社会影响因素决定其判断调整。

Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning

论文配图:Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning
图 1 · 摘自论文原文
  • 通过三类社会影响维度(距离、来源、群体)建模判断修正过程
  • 模型更易接受相近观点,受自称原有观点者影响更大,对群体压力反应不同
  • 为区分合理改变认知与盲目迎合提供可操作框架,适合对齐研究者参考

构建具备社会适应能力的大语言模型,不仅需减少单维的盲从倾向,还需学会在采纳他人观点与坚守道德判断之间做出区分。我们研究了这一权衡背后的抵抗-顺从机制。三个实验表明,模型的判断修正呈现三个平行于人类社会心理学的经典维度:外部观点与模型初始立场的距离、观点来源归属、支持该观点的群体结构。模型对邻近观点更开放,更易被归因于自身先前判断的观点影响,且对群体压力的响应存在差异。这些发现将盲从重新定义为一种由社会影响塑造的更广泛判断更新过程。本框架为区分建设性认知修正与盲目顺从提供了原则基础,有助于在涉及道德后果的交互中实现更好的对齐。

原文摘要 · Abstract (English)

Building socially calibrated large language models, which can learn from others without simply yielding to them, requires more than reducing sycophancy as a one-dimensional failure mode. Models must distinguish when to incorporate others' perspectives from when to maintain a well-grounded moral judgment. We study the broader resistance-compliance process governing this distinction. Across three studies, we show that models' judgment revision is structured along three dimensions that parallel classic phenomena in human social psychology: the distance between an incoming view and the model's initial position, the source attribution of that view, and the coalition structure supporting it. Models are generally more receptive to nearby positions, more influenced by views presented as their own prior judgments, and differently responsive to group pressure. These findings recast sycophancy as one expression of a broader judgment-updating process shaped by social influence. Our framework provides a principled basis for distinguishing constructive belief revision from sycophantic compliance, thereby supporting better alignment in morally consequential interactions.

大模型对齐道德推理社会影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。