让多个AI辩论时记住过往经验,自动调整信任度,防止集体犯错。
Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation

- 用历史辩论记录动态修正每个AI的先验认知
- 根据过往表现给参与方分配可信权重,减少错误影响
- 适合需要多智能体协作推理的复杂任务场景
多智能体辩论(MAD)通过多个智能体迭代讨论提升大模型的推理能力。然而,当多数智能体最初形成错误共识时,辩论过程反而会放大错误,存在‘共性误解’问题。现有方法仅关注同伴偏差,未解决智能体固有的概念先验偏见。为此,本文提出R²-MAD框架,为智能体配备从过往辩论中积累的经验记忆。该框架通过双重机制干预两种失效模式:基于当前共识水平,采用辩论状态感知的检索策略动态校准概念先验;随后利用检索到的历史经验评估各智能体可靠性,生成置信权重以调节同伴影响力。在多个基准测试上的实验表明,R²-MAD在单智能体及现有MAD基线之上均实现稳定性能提升。
原文摘要 · Abstract (English)
Multi-agent debate (MAD) improves the reasoning capabilities of large language models by having multiple agents iteratively refine their responses through discussion. However, MAD suffers from a critical vulnerability known as shared misconception: when a majority of agents initially converge on an incorrect answer, the debate process tends to amplify rather than correct the error. Existing methods primarily address peer skew but leave the agents' inherently biased concept priors unaddressed. To mitigate this systematic weakness, we propose R$^2$-MAD (Remember and Reweight for Multi-Agent Debate), a framework that equips agents with an experience memory accumulated from past debates. R$^2$-MAD intervenes on both failure modes through two complementary mechanisms: A debate-state-aware retrieval policy dynamically calibrates the concept prior by retrieving relevant historical evidence based on the current consensus level. Then these retrieved experiences provide a basis for estimating per-agent reliability, yielding confidence weights to modulate peer influence. Experiments on various benchmarks show that R$^2$-MAD achieves consistent improvements over existing single-agent and MAD baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。