arXiv:2502.18439cs.AI2025-02ACL被引 89

用强化学习训练多个大模型协同讨论,提升问答准确性。

MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning

  • 多模型独立生成后通过多轮讨论协作改进答案。
  • 验证器评分驱动强化学习,使协作正确率提升23%以上。
  • 适合研究多智能体协作与大模型后训练的读者。

利用多个大语言模型构建协同多智能体流程展现出巨大潜力。然而,以往研究多依赖预训练模型的固有能力进行协作,近期研究表明这未必能有效提升性能。本文提出一种新型后训练范式MAPoRL(基于强化学习的多智能体协同后训练),旨在显式激发协同行为并释放多智能体框架的潜力。在MAPoRL中,多个大模型首先独立生成响应,并通过多轮讨论共同优化最终答案。最后,一个验证器评估答案和讨论质量,根据正确性打分并激励修正与说服性对话。该分数作为协同训练奖励,通过多智能体强化学习最大化。不同于现有后训练方法,MAPoRL倡导多模型联合使用强化学习进行协同训练,以提升泛化能力。分析与实验表明,单独训练各模型无法有效诱导协作,而多智能体协同训练显著提升跨基准测试的表现,并可泛化至未见领域。

原文摘要 · Abstract (English)

Leveraging multiple large language models (LLMs) to build collaborative multi-agentic workflows has demonstrated significant potential. However, most previous studies focus on prompting the out-of-the-box LLMs, relying on their innate capability for collaboration, which may not improve LLMs' performance as shown recently. In this paper, we introduce a new post-training paradigm MAPoRL (Multi-Agent Post-co-training for collaborative LLMs with Reinforcement Learning), to explicitly elicit the collaborative behaviors and further unleash the power of multi-agentic LLM frameworks. In MAPoRL, multiple LLMs first generate their own responses independently and engage in a multi-turn discussion to collaboratively improve the final answer. In the end, a MAPoRL verifier evaluates both the answer and the discussion, by assigning a score that verifies the correctness of the answer, while adding incentives to encourage corrective and persuasive discussions. The score serves as the co-training reward, and is then maximized through multi-agent RL. Unlike existing LLM post-training paradigms, MAPoRL advocates the co-training of multiple LLMs together using RL for better generalization. Accompanied by analytical insights, our experiments demonstrate that training individual LLMs alone is insufficient to induce effective collaboration. In contrast, multi-agent co-training can boost the collaboration performance across benchmarks, with generalization to unseen domains.

多智能体强化学习协同问答大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。