arXiv:2606.01637cs.CLcs.AI2026-06被引 1

研究发现模型易被同伴误导,纠错能力远低于引入错误。

Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity

论文配图:Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity
图 1 · 摘自论文原文
  • 设计控制实验,测试共识与权威标签对模型修订的影响
  • 正确模型被误导概率是纠正错误的3倍以上
  • 思维链等干预方法无法同时减少有害修订和保留有益修订

大型语言模型在多智能体系统中频繁响应其他智能体的回答,存在认知从众风险:模型可能因他人意见一致而放弃自身判断。本文通过受控实验,让模型先独立回答问题,再观察模拟同伴回应后做出最终决定。我们操控了两种社会线索——共识结构和同伴权威标签,评估其对有益与有害修订的影响。在四个开源大模型与七个问答数据集上,结果表明:同伴一致意见使原本正确的模型更易被误导,而纠正错误的能力显著较弱;权威标签会显著提高模型采纳对应答案的概率,无论答案正确与否。更令人担忧的是,常见的推理干预(如思维链、反思)无法可靠降低有害修订,同时保持有益修订。研究建议,多智能体系统应验证同伴答案,而非简单聚合。

原文摘要 · Abstract (English)

Large language models are increasingly used in multi-agent systems, where they see and respond to other agents' answers. A key risk is conformity: a model may abandon its own answer simply because others agree on a different one. Prior studies show that LLMs often revise toward a majority answer, but it remains unclear whether these revisions help correct mistakes as often as they introduce new errors. In this paper, we conduct a controlled study in which an LLM first answers a question, then sees simulated peer responses before making a final decision. We manipulate two social cues: consensus structure and authority labels assigned to peers, and measure how they influence beneficial and harmful revisions. Across four open-weight LLMs and seven QA datasets, we find that peer agreement makes it much easier to mislead initially correct models than to correct initially wrong ones. Authority labels make models more likely to choose the endorsed answer, regardless of whether it is correct. More concerningly, generic reasoning interventions such as chain-of-thought and reflection do not reliably reduce harmful revision while preserving beneficial revision. These findings suggest that multi-agent LLM systems should verify peer answers rather than simply aggregate them.

大模型从众偏差多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。