通过共享同伴谄媚程度评分,显著降低大模型在协作讨论中的盲从行为。
Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems

- 给多智能体系统引入同伴谄媚度评分机制
- 使讨论准确率提升10.5%,有效抑制错误传播
- 适合关注模型协作可信度的研究者
大型语言模型常表现出谄媚倾向:即使与自身观点冲突,也倾向于附和用户立场。现有研究多集中于单智能体场景,而多智能体协同中的此类现象仍待探索。本文探究了智能体对同伴谄媚程度的认知是否影响讨论结果。我们使用六种开源LLM进行受控实验,为智能体提供基于静态(讨论前)和动态(在线)策略计算的同伴谄媚度评分。结果显示,引入谄媚先验可降低谄媚型同伴的影响,缓解错误级联,并使最终讨论准确率提升绝对10.5%。该方法轻量高效,能有效抑制模型在讨论中的谄媚行为,进而提升下游任务准确性。
原文摘要 · Abstract (English)
Large language models (LLMs) often exhibit sycophancy: agreement with user stance even when it conflicts with the model's opinion. While prior work has mostly studied this in single-agent settings, it remains underexplored in collaborative multi-agent systems. We ask whether awareness of other agents' sycophancy levels influences discussion outcomes. To investigate this, we run controlled experiments with six open-source LLMs, providing agents with peer sycophancy rankings that estimate each peer's tendency toward sycophancy. These rankings are based on scores calculated using various static (pre-discussion) and dynamic (online) strategies. We find that providing sycophancy priors reduces the influence of sycophancy-prone peers, mitigates error-cascades, and improves final discussion accuracy by an absolute 10.5%. Thus, this is a lightweight and efficient way to reduce model sycophancy during discussions and subsequently improve downstream accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。