检验多智能体辩论是否真能提升答案质量,发现表面分歧多为表象。
A Layered Analysis of Disagreement And Answer Quality in Multi-Agent LLM Debate
- 通过四类测量分析辩论中真实分歧与表态一致性
- 不同语气下同意率相差超50个百分点,但最终答案无质量提升
- 反驳行为多受指令影响,且多数不反映真实立场持久性
多智能体辩论被普遍认为可通过暴露真实分歧来提升答案质量,但这一机制极少被验证。本文引入四项测量:(A) 辩论者报告的同意度;(B) 回复文本是否真正提出反驳;(C) 撤除诱发指令后立场是否持续;(D) 对开放权重模型,分析其自身令牌概率中的立场倾向。在750场针对GlobalOpinionQA的辩论中,评估三种模型委员会在友好、中立、敌对三种语气下的表现。结果显示:(A) 语气显著影响报告同意度,友好与敌对间差异达50.4个百分点;(B) 仅读回复文本即可复现相同模式;(C) 分歧部分源于诱发指令:移除敌对指令后,立场回归一致的幅度比保留指令时高出23.1点,首轮独立回溯仅11/28出现在回复中,整体分析具有统计显著性(p=0.016);(D) 反对意见更削弱立场强度而非改变方向。最终答案未见质量提升:经偏差校验的评审团判定299/299持平(排除明显差异),可控任务准确率未变,无偏差检查的评审团曾以66%胜率认为辩论更优——实为阅读顺序导致的假象。总体表明,LLM辩论虽可改变言论表达,但对持久立场或最终答案质量的改进证据薄弱。
原文摘要 · Abstract (English)
Multi-agent debate, in which several LLMs exchange arguments before answering, is widely assumed to improve answer quality by surfacing genuine disagreement. That mechanism is rarely checked. We introduce four measurements: (A) the agreement a debater reports; (B) whether its reply text actually pushes back; (C) whether the position persists once the eliciting instruction is removed; and (D) for open-weight models, the stance response in the debater's own token log-probabilities. We evaluate three-model committees debating open-ended GlobalOpinionQA across 750 debates under three tones: friendly (seek common ground), neutral, and hostile (stress-test every position). (A) Tone strongly reshapes reported agreement: full agreement differs by 50.4 percentage points between the friendly and hostile endpoints. (B) A judge that reads only the reply text, never the self-report or the condition, recovers the same pattern. (C) The dissent appears partly tied to the instruction that elicited it: labels revert toward agreement 23.1 points more often after deleting the hostile instruction than under a matched re-ask that keeps it; question-weighted inference is inconclusive on first-round turns alone (p=0.0625), significant pooling all rounds (p=0.016), and only 11/28 first-round reversions also appear in the reply text. (D) Opposing arguments weaken a debater's stance margin more consistently than they shift its direction. For final answers we detect no quality gain: a bias-checked jury returns 299/299 ties (ruling out only large differences), accuracy on a verifiable control task is unchanged, and a jury without the bias check had declared debate the winner 66% of the time -- an artifact of reading order. Taken together, LLM debate readily changes what agents say, but we find much weaker evidence that it changes what they persistently endorse or improves the quality of the final answer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。