用自对弈训练大模型在合作与对抗中提升多智能体推理能力。
MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
- 通过回合级优势估计和个性化归一化,解决多智能体长程信用分配难题。
- 在未见游戏中性能提升最高达28.7%,跨任务基准平均提升3.5%。
- 适用于需要策略性协作或竞争的智能体系统,如复杂推理场景。
构建能在多智能体系统(MAS)中有效协作与竞争的大语言模型(LLM),是迈向更高级智能的关键一步。尽管强化学习(RL)在单智能体推理任务中表现优异,但其扩展至多轮、多智能体场景仍面临长程信用分配与个体优势估计的挑战。为此,我们提出MARSHAL,一种端到端的强化学习框架,通过战略型大模型的自对弈,在合作与对抗游戏中激励多智能体推理。MARSHAL采用回合级优势估计器,使学习信号与每轮交互对齐以实现精准信用分配;同时引入智能体特定的优势归一化机制,稳定多智能体训练过程。基于Qwen3-4B训练的MARSHAL智能体,在未见游戏中性能最高提升28.7%。更重要的是,自对弈获得的能力具有泛化性,在推理基准测试中表现出一致提升:集成至主流多智能体系统后,于AIME上零样本性能提升10.0%,于GPQA-Diamond上提升7.6%,所有基准平均提升3.5%。这些结果确立了战略游戏中的自对弈作为发展可泛化多智能体推理能力的有效路径。
原文摘要 · Abstract (English)
Developing Large Language Models (LLMs) to cooperate and compete effectively within multi-agent systems (MASs) is a critical step towards more advanced intelligence. While reinforcement learning (RL) has proven effective for enhancing reasoning in single-agent tasks, its extension to multi-turn, multi-agent scenarios remains underexplored due to the challenges of long-horizon credit assignment and agent-specific advantage estimation. To address these challenges, we introduce MARSHAL, an end-to-end RL framework that incentivizes Multi-Agent Reasoning through Self-play witH strAtegic LLMs in both cooperative and competitive games. MARSHAL features a turn-level advantage estimator that aligns learning signals with each interaction for credit assignment, and an agent-specific advantage normalization to stabilize multi-agent training. By learning with self-play across cooperative and competitive games, MARSHAL agents trained from Qwen3-4B develop strong strategic abilities, with up to 28.7% performance improvements in held-out games. More importantly, the capability acquired through self-play generalizes beyond games, yielding consistent performance gains of MASs in reasoning benchmarks. When integrated into leading MASs, our MARSHAL agent achieves significant zero-shot performance gains of up to 10.0% on AIME, 7.6% on GPQA-Diamond, and 3.5% on average across all benchmarks. These results establish self-play in strategic games as a powerful approach for developing generalizable multi-agent reasoning capabilities in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。