arXiv:2606.00820cs.CL2026-06被引 1

拆解大模型辩论中立场趋同的三种机制,发现多数趋同实为盲目从众。

Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate

  • 提出三源分解框架,区分自发波动、受立场影响的从众和理性说服。
  • 37%的判断在自我反思后改变,严格从众率达29%,且多导致错误。
  • 可基于初始特征预测有害从众,干预后错误率下降13.6个百分点。

多智能体辩论(MAD)是提升大模型推理能力的有前景策略,但当智能体趋于一致时,这种收敛究竟是真实讨论的结果,还是社会性服从?我们发现,传统“答案翻转率”混淆了三种不同机制:自发不稳定性、立场诱导的从众以及推理引发的说服。通过控制反事实条件,我们的三源分解框架将它们分离。在主实验集MMLU-Pro中,37%的智能体-问题观测在仅自我反思下发生改变;稳健性测试显示,GPQA-Diamond及三个模型家族存在显著模型依赖的不稳定性;严格从众在主设置中占29%,且在模型复现中持续有害(正确转错误占比57%-77%)。信息梯度控制实验表明,即使无效推理也导致20%-39%的顽固智能体采纳错误,且具有类似推理形式的表达具有强说服力。有害从众可由第0轮特征预测(AUC=0.79),针对性干预使其降低13.6个百分点(p<0.001)。然而,若无正确性标签或自我反思控制,减少同伴采纳并不能提升准确率,因有害与有益影响无法区分。

原文摘要 · Abstract (English)

Multi-agent debate (MAD) is a promising strategy for improving LLM reasoning, but when agents converge on a shared answer, it is unclear whether that convergence reflects genuine deliberation or social compliance. We show that the conventional answer flip rate conflates three distinct mechanisms: spontaneous instability, stance-induced conformity, and reasoning-induced persuasion. Our three-source decomposition framework isolates each through controlled counterfactual conditions. In the primary MMLU-Pro setting, 37% of agent-question observations change under self-reflection alone, while robustness tests show substantial model-dependent instability across GPQA-Diamond and three model families; strict conformity is 29% in the primary setting and remains predominantly harmful across model replications (57-77% correct-to-wrong). A controlled information-gradient experiment reveals that even vacuous reasoning is associated with 20-39% error adoption among resistant agents, with reasoning-like presentation carrying substantial persuasive weight. Harmful conformity can be predicted from Round 0 features (AUC = 0.79), and risk-targeted intervention reduces it by 13.6 percentage points (p < 0.001). However, without correctness labels or self-reflection controls, reducing peer adoption does not improve accuracy, because harmful and beneficial influence cannot be distinguished.

大模型辩论从众行为推理评估机制分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。