arXiv:2603.21574cs.AI2026-03被引 4

提出新框架提升多智能体推理的稳定性与可解释性。

Adaptive Robust Estimator for Multi-Agent Reinforcement Learning

  • 分三阶段生成、批评、重写,明确各智能体贡献
  • 在噪声奖励下仍保持性能稳定,优于基线方法
  • 适合需要可靠协作的复杂推理任务

多智能体协作已成为增强大语言模型推理能力的有效范式,但其交互层面存在生成、批评与修订模糊的问题,导致智能体间责任分配困难。此外,该场景下的策略优化易受长尾且嘈杂奖励的影响,偏差优势估计并引发训练不稳定甚至发散。为此,我们提出一种稳健的多智能体强化学习框架,包含双智能体答案-批评-重写(DACR)与自适应鲁棒估计器(ARE)。DACR将推理分解为结构化三阶段流程,实现对各智能体边际贡献的显式归因;ARE在多智能体策略优化中提供批经验均值的鲁棒估计。在数学推理与具身智能基准测试中,即使在噪声奖励条件下,该方法在同质与异质设置下均持续优于基线。结果表明其对奖励噪声具有更强鲁棒性,训练动态更稳定,有效防止由噪声奖励信号引起的优化失败。

原文摘要 · Abstract (English)

Multi-agent collaboration has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models, yet it suffers from interaction-level ambiguity that blurs generation, critique, and revision, making credit assignment across agents difficult. Moreover, policy optimization in this setting is vulnerable to heavy-tailed and noisy rewards, which can bias advantage estimation and trigger unstable or even divergent training. To address both issues, we propose a robust multi-agent reinforcement learning framework for collaborative reasoning, consisting of two components: Dual-Agent Answer-Critique-Rewrite (DACR) and an Adaptive Robust Estimator (ARE). DACR decomposes reasoning into a structured three-stage pipeline: answer, critique, and rewrite, while enabling explicit attribution of each agent's marginal contribution to its partner's performance. ARE provides robust estimation of batch experience means during multi-agent policy optimization. Across mathematical reasoning and embodied intelligence benchmarks, even under noisy rewards, our method consistently outperforms the baseline in both homogeneous and heterogeneous settings. These results indicate stronger robustness to reward noise and more stable training dynamics, effectively preventing optimization failures caused by noisy reward signals.

多智能体强化学习鲁棒性协作推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。