研究大模型在多人协作中如何受同伴影响,发现大模型更抗干扰。
LLMs Can't Handle Peer Pressure: Crumbling under Multi-Agent Social Interactions
- 构建可控互动的答题基准测试KAIROS,模拟真实社交环境。
- 大模型比小模型更抗社会压力,提示词可有效缓解小模型弱点。
- 仅精心设计的强化学习训练能让小模型稳定提升抗干扰能力。
大型语言模型(LLMs)正越来越多地被集成到多智能体系统(MAS)中,同伴互动会影响个体决策。以往研究主要关注从众偏差,本文扩展视角,考察LLMs如何基于过往互动建立关系、识别并整合高质量同伴信息,以及抵抗误导性输入——这些能力对复杂社交动态下的集体智能至关重要。我们提出KAIROS基准,通过可精确控制历史互动与当前轮次中同伴关系水平和行为的问答式协作场景,实现对关系、同伴行为与模型自信心共同影响决策的系统分析。利用KAIROS,我们评估了提示工程、监督微调及基于组相对策略优化(GRPO)的强化学习。结果表明,模型规模是调节社会影响力敏感度的关键因素:大模型更具韧性,且能通过提示缓解;小模型仍易受影响。只有经过仔细配置的GRPO训练才能为小模型带来一致的鲁棒性与性能提升。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly integrated into multi-agent systems (MAS), where peer interactions shape individual decisions. While prior work has mainly examined conformity bias, we broaden the view to include how LLMs build rapport from prior interactions, discern and integrate high-quality peer information, and resist misleading inputs-abilities essential for achieving collective intelligence under complex social dynamics. We introduce KAIROS, a benchmark that simulates quiz-style collaboration with peer agents whose rapport levels and behaviours can be precisely controlled in both historical interactions and the current round. This unified setup enables systematic analysis of how rapport, peer actions, and the model's self-confidence jointly influence decision-making. Using KAIROS, we evaluate prompting, supervised fine-tuning, and reinforcement learning via Group Relative Policy Optimisation (GRPO). Results show that model scale is a primary factor moderating susceptibility to social influence: larger models are more resilient and benefit from prompting-based mitigation, whereas smaller models remain vulnerable. Only carefully configured GRPO training yields consistent robustness and performance gains for small models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。