arXiv:2601.09667cs.AIcs.CL2026-01被引 3

多智能体推理时动态协作,提升复杂任务准确率

Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning

  • 推理阶段引入结构化文本经验,多专家协作讨论
  • 医学、数学等任务上平均提升3.67%,最高超单体模型8.67%
  • 无需调参,对分布外数据鲁棒,适合高可靠性场景

多智能体系统已发展为许多应用中由大语言模型驱动的协作伙伴,其鲁棒性得益于多样性与交叉验证。然而,多智能体强化学习(MARL)训练资源消耗大且不稳定:队友协同适应导致非平稳性,奖励常稀疏且方差高。为此,我们提出多智能体推理时强化学习(MATTRL),在推理阶段注入结构化文本经验,形成多专家团队进行多轮讨论,检索并整合测试时经验,最终达成共识决策。我们研究了信用分配机制以构建逐轮经验池,并将其重新注入对话。在医学、数学和教育等挑战性基准上,MATTRL相较多智能体基线平均提升3.67%,相比同类单智能体基线提升8.67%。消融实验分析了不同信用分配方案对训练结果的影响。MATTRL提供了一条稳定、高效且无需调参的多智能体推理路径,具备分布外鲁棒性。

原文摘要 · Abstract (English)

Multi-agent systems have evolved into practical LLM-driven collaborators for many applications, gaining robustness from diversity and cross-checking. However, multi-agent RL (MARL) training is resource-intensive and unstable: co-adapting teammates induce non-stationarity, and rewards are often sparse and high-variance. Therefore, we introduce \textbf{Multi-Agent Test-Time Reinforcement Learning (MATTRL)}, a framework that injects structured textual experience into multi-agent deliberation at inference time. MATTRL forms a multi-expert team of specialists for multi-turn discussions, retrieves and integrates test-time experiences, and reaches consensus for final decision-making. We also study credit assignment for constructing a turn-level experience pool, then reinjecting it into the dialogue. Across challenging benchmarks in medicine, math, and education, MATTRL improves accuracy by an average of 3.67\% over a multi-agent baseline, and by 8.67\% over comparable single-agent baselines. Ablation studies examine different credit-assignment schemes and provide a detailed comparison of how they affect training outcomes. MATTRL offers a stable, effective and efficient path to distribution-shift-robust multi-agent reasoning without tuning.

多智能体推理增强强化学习LLM协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。