同质大模型辩论反而易出错,自修正更高效可靠。
The Cost of Consensus: Isolated Self-Correction Prevails Over Unguided Homogeneous Multi-Agent Debate
- 用10个同规模模型反复辩论,对比单独修正和随机干扰效果。
- 辩论失败率高达70%,多数模型盲目跟风,正确答案常被投票淘汰。
- 单独自我修正比集体辩论省3倍以上算力,准确率还更高。
多智能体辩论中,多个大模型通过迭代推理与投票选出答案,常被认为能抑制幻觉。然而,同质化辩论的失效机制仍不清晰。本研究在GSM-Hard和MMLU-Hard两个高难度基准上,对10个同规模模型(Qwen2.5-7B、Llama-3.1-8B、Ministral-3-8B)进行了三轮辩论实验。对比了同伴辩论、孤立自修正及引入无关问题推理的随机噪声控制。结果揭示三种失败路径:谄媚顺从(采纳多数意见最高达85.5%)、上下文脆弱性(同伴推理使原有正确推理崩溃,脆弱率高达70.0%)、共识崩溃(多数投票排除已存在于生成池中的正确答案,差距最高达32.3个百分点)。在通信密度K∈{2,4,9}和采样温度T∈{0.4,0.7}的消融实验中发现,低暴露(K=2)即引发高度顺从,初始多样性越高,顺从越强。所有配置下,辩论耗时2.1–3.4倍于自修正,最多达28,631词元/题,且准确率更低或相当。结论表明,在7–8B参数量级下,无结构角色的同质团队无法从无引导的同伴交互中获益,孤立自修正始终提供更优的成本-精度权衡。
原文摘要 · Abstract (English)
Multi-agent debate, where teams of LLMs iteratively exchange rationales and vote on answers, is widely deployed under the assumption that peer review filters hallucinations. Yet the failure dynamics of homogeneous debate remain poorly understood, therefore we report findings from a controlled empirical study of teams of $N{=}10$ homogeneous agents (Qwen2.5-7B, Llama-3.1-8B, Ministral-3-8B) across $R{=}3$ debate rounds on two high-difficulty benchmarks (GSM-Hard and MMLU-Hard). We compare peer debate against isolated self-correction and a stochastic noise control that injects rationales from unrelated problems. We decompose debate failure into three model-dependent pathways: sycophantic conformity, where agents uncritically adopt majority answers (modal adoption up to 85.5%); contextual fragility, where peer rationales destabilize previously correct reasoning (vulnerability rate up to 70.0%); and consensus collapse, where plurality voting discards correct answers already present in the generation pool (oracle gap up to 32.3 percentage points). Ablations over communication density ($K \in \{2,4,9\}$) and sampling temperature ($T \in \{0.4, 0.7\}$) show that conformity reaches high levels at minimal peer exposure ($K{=}2$) and intensifies with greater initial diversity. Across all configurations, debate consumes 2.1-3.4$\times$ more tokens (up to 28,631 tokens per problem) than self-correction for equal or lower accuracy. Our results indicate that, within the 7-8B parameter class, homogeneous teams without structured roles do not benefit from unguided peer exchange, and that isolated self-correction consistently offers a more favorable cost-accuracy tradeoff.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。