arXiv:2509.14680cs.MAcs.LG2025-09被引 8

用大模型生成专家示范,让多智能体高效协作。

LEED: A Highly Efficient and Scalable LLM-Empowered Expert Demonstrations Framework for Multi-Agent Reinforcement Learning

  • 用大模型自动生成环境交互指令,生成高质量示范数据
  • 各智能体基于示范数据优化本地策略,实现高效个性化学习
  • 在复杂场景下显著提升训练效率和可扩展性,适合大规模多智能体系统

多智能体强化学习(MARL)在复杂环境中的智能决策中具有巨大潜力,但随着智能体数量增加,协调与可扩展性面临瓶颈。为此,我们提出面向多智能体强化学习的大型语言模型赋能专家示范框架(LEED)。LEED包含示范生成(DG)模块和策略优化(PO)模块。DG模块利用大语言模型生成环境交互指令,从而生成高质量示范数据;PO模块采用去中心化训练范式,每个智能体使用生成的示范数据构建专家策略损失,并与自身策略损失结合,实现基于专家知识与个体经验的本地策略有效个性化优化。实验表明,相比最先进基线,LEED在样本效率、时间效率和鲁棒可扩展性方面均表现更优。

原文摘要 · Abstract (English)

Multi-agent reinforcement learning (MARL) holds substantial promise for intelligent decision-making in complex environments. However, it suffers from a coordination and scalability bottleneck as the number of agents increases. To address these issues, we propose the LLM-empowered expert demonstrations framework for multi-agent reinforcement learning (LEED). LEED consists of two components: a demonstration generation (DG) module and a policy optimization (PO) module. Specifically, the DG module leverages large language models to generate instructions for interacting with the environment, thereby producing high-quality demonstrations. The PO module adopts a decentralized training paradigm, where each agent utilizes the generated demonstrations to construct an expert policy loss, which is then integrated with its own policy loss. This enables each agent to effectively personalize and optimize its local policy based on both expert knowledge and individual experience. Experimental results show that LEED achieves superior sample efficiency, time efficiency, and robust scalability compared to state-of-the-art baselines.

多智能体强化学习大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。