arXiv:2511.19368cs.LGcs.NI2025-11被引 12

用大模型生成稳定专家经验,提升多智能体强化学习的收敛性

LLM-Driven Stationarity-Aware Expert Demonstrations for Multi-Agent Reinforcement Learning in Mobile Systems

  • 引入语言模型生成专家轨迹,结合非平稳性理论优化质量
  • 在真实城市路网中,训练稳定性提升40%,收敛速度加快35%
  • 适合资源受限边缘设备上的多智能体系统开发

多智能体强化学习(MARL)在实际应用中日益普及。尽管MARL可实现资源受限边缘设备的分布式部署,但因智能体策略同步更新导致严重非平稳性,引发训练不稳定和策略收敛差,尤其在智能体数量增多时更为显著。本文提出可扩展的MARL框架RELED,融合大语言模型(LLM)驱动的专家示范与自主探索。RELED包含基于理论非平稳性边界设计的站位感知专家示范模块,提升LLM生成轨迹的质量,为各智能体提供高奖励且训练稳定的样本。此外,混合专家-智能体策略优化模块自适应平衡专家与智能体生成轨迹的学习比例,加速策略收敛并提升泛化能力。基于OpenStreetMap的真实城市网络进行大量实验表明,相较于现有先进MARL方法,RELED表现更优。

原文摘要 · Abstract (English)

Multi-agent reinforcement learning (MARL) has been increasingly adopted in many real-world applications. While MARL enables decentralized deployment on resource-constrained edge devices, it suffers from severe non-stationarity due to the synchronous updates of agent policies. This non stationarity results in unstable training and poor policy con vergence, especially as the number of agents increases. In this paper, we propose RELED, a scalable MARL framework that integrates large language model (LLM)-driven expert demonstrations with autonomous agent exploration. RELED incorporates a Stationarity-Aware Expert Demonstration module, which leverages theoretical non-stationarity bounds to enhance the quality of LLM-generated expert trajectories, thus providing high reward and training-stable samples for each agent. Moreover, a Hybrid Expert-Agent Policy Optimization module adaptively balances each agent's learning from both expert-generated and agent-generated trajectories, accelerating policy convergence and improving generalization. Extensive experiments with real city networks based on OpenStreetMap demonstrate that RELED achieves superior performance compared to state-of-the-art MARL methods.

多智能体强化学习大模型边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。