arXiv:2606.13197cs.AI2026-06

让AI辩论更智能:按需启动、自动停止、过滤异常,提升推理准确率。

ARMOR-MAD: Adaptive Routing for Heterogeneous Multi-Agent Debate in Large Language Model Reasoning

论文配图:ARMOR-MAD: Adaptive Routing for Heterogeneous Multi-Agent Debate in Large Language Model Reasoning
图 1 · 摘自论文原文
  • 根据答案一致性决定是否辩论,节省计算资源。
  • 辩论提前收敛时自动停止,效率提升明显。
  • 识别并弱化异常答案,提高最终决策质量。

多智能体辩论(MAD)可提升大模型推理能力,但固定辩论流程常导致计算浪费,并放大相似智能体的共现错误。本文提出无需训练的异构辩论框架ARMOR-MAD,将辩论视为条件计算。该框架包含三个组件:预辩论一致性路由(PAR)判断初始回答是否需要辩论;早期共识终止评估器(EASE)在达成共识后提前停止辩论;语义离群检测(SOD)在聚合时降低异常答案权重。在MATH Level 5、GSM8K、MMLU和MMLU-Pro上,ARMOR-MAD均优于相同模型池的固定轮次异构辩论,准确率分别达到65.5%、96.5%、90.0%和81.5%。结果表明,真实模型异质性与基于一致性的控制对提升MAD的准确性与效率至关重要。

原文摘要 · Abstract (English)

Multi-agent debate (MAD) can improve large language model reasoning, but fixed debate pipelines often waste computation and can amplify correlated errors among similar agents. We propose ARMOR-MAD, a training-free heterogeneous MAD framework that treats debate as conditional computation. ARMOR-MAD combines three components: Pre-debate Agreement Routing (PAR) decides whether independently generated Round-0 answers require debate; Early Agreement Stopping Evaluator (EASE) stops debate after convergence; and Semantic Outlier Detection (SOD) down-weights abnormal final answers during aggregation. Across MATH Level 5, GSM8K, MMLU, and MMLU-Pro, ARMOR-MAD consistently improves over fixed-round heterogeneous debate with the same model pool, reaching 65.5\%, 96.5\%, 90.0\%, and 81.5\% accuracy, respectively. The results suggest that genuine model heterogeneity and agreement-based control are both important for making MAD more accurate and efficient.

多智能体推理优化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。