arXiv:2503.16814cs.LGcs.CL2025-03Transactions of th…被引 2

解决大模型推理中因观点固化导致的错误强化问题

From Belief Entrenchment to Robust Reasoning in LLM Agents

  • 通过引入先验知识和视角多样性重构辩论机制
  • 在新基准上比ReAct提升9.5%准确率,胜率高出19%
  • 适合需要高可靠推理的复杂任务场景

多智能体辩论(MAD)是提升大语言模型推理能力的有前景方法,但常因观点固化而加剧共同错误。我们发现该问题源于两个根源:一是模型静态初始信念存在偏差,二是辩论动态趋同放大多数意见。为此提出DReaMAD框架,先通过策略性先验知识提取修正初始信念,再通过强制视角多样性重塑辩论过程。在新构建的MetaNIM Arena基准上验证,该方法显著缓解了固化现象,相比ReAct提示提升9.5%准确率,标准MAD的胜率高出19.0%。

原文摘要 · Abstract (English)

Multi-Agent Debate (MAD) has emerged as a promising inference scaling method for Large Language Model (LLM) reasoning. However, it frequently suffers from belief entrenchment, where agents reinforce shared errors rather than correcting them. Going beyond merely identifying this failure, we decompose it into two distinct root causes: (1) the model's biased $\textit{static initial belief}$ and (2) $\textit{homogenized debate dynamics}$ that amplify the majority view regardless of correctness. To address these sequentially, we propose $\textbf{DReaMAD}$ $($$\textbf{D}$iverse $\textbf{Rea}$soning via $\textbf{M}$ulti-$\textbf{A}$gent $\textbf{D}$ebate with Refined Prompt$)$. Our framework first rectifies the static belief via strategic prior knowledge elicitation, then reshapes the debate dynamics by enforcing perspective diversity. Validated on our new $\textit{MetaNIM Arena}$ benchmark, $\textbf{DReaMAD}$ significantly mitigates entrenchment, achieving a +9.5\% accuracy gain over ReAct prompting and a +19.0\% higher win rate than standard MAD.

大模型推理多智能体辩论机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。