arXiv:2608.11247cs.AIcs.MA2026-08

大模型在集体讨论中易被错误多数意见误导,本文揭示其抗干扰与学习能力存在权衡边界。

Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier

论文配图:Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier
图 1 · 摘自论文原文
  • 提出抵抗性与接受性双维度评估模型在群体协作中的表现
  • 错误一致多数可使模型正确率下降最高71%,84-89%的错误答案与同伴一致
  • 发现唯一能同时提升两种能力的方法是自我推理,适合需要高可靠性的决策场景

近期语言模型发展使多模型协作成为可能,各模型在回答前可观察他人输出,导致群体意见与自身知识竞争。当多数意见错误时,可能颠覆模型原本正确的判断。我们在23个开源模型、19种条件和3个数据集上测试,生成超百万条评分响应。在统一错误多数下,模型正确答案被逆转比例达:MMLU 22.8%,GPQA 54.8%,SimpleQA 71.0%;其中84-89%的被逆转答案与同伴一致。现有缓解策略仅提升抵抗性(保持正确答案的能力),但忽视接受性(纠正自身错误并采纳正确同伴答案的能力)。我们评估六种方法在两维上的表现,发现所有方法均以牺牲一方换取另一方,且结果落在单一抵抗性-接受性前沿线上,决定系数R²为0.80至0.90之间。反思(Reflection)是最强方法,在MMLU上提升7.9点抵抗性,但损失15.3点接受性。推理(Reasoning)是唯一例外:在可自主推导的MMLU题目上,同时提升7.2点抵抗性和9.6点接受性,是唯一能兼顾两者的干预方式。

原文摘要 · Abstract (English)

Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iteratively improving, transforming, and extending each other's outputs. Each agent sees what the others assert before it answers, so peer opinion competes with the model's own parametric knowledge, and a wrong majority can overturn an answer the model would otherwise get right. We measure that displacement in 23 open-weight models, 19 conditions, and three datasets, yielding more than a million graded responses. A unanimous wrong majority reverses 22.8% of a model's correct MMLU answers, 54.8% on GPQA and 71.0% on SimpleQA, and 84-89% of the reversed answers match the peers' answers. Existing mitigations aim to increase Resistance, the rate at which a model keeps its correct answer under this pressure, which is only half of what a collaborating agent needs. We pair it with Receptivity, the rate at which a model adopts a correct peer answer after initially answering incorrectly. We score six methods on both axes, four drawn from prior work and two of our own. Each gains Resistance only by losing Receptivity, and their means fall on a single Resistance-Receptivity frontier with $R^2$ between 0.80 and 0.90. Reflection, the strongest published method, gains 7.9 points of MMLU Resistance and gives up 15.3 of Receptivity. Reasoning is the one exception. On GPQA and SimpleQA it trades like the rest, but on the MMLU subjects whose answers a model can derive for itself it raises Resistance by 7.2 points and Receptivity by 9.6 at once, the only intervention we find that improves both.

大模型协作群体偏差推理机制评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。