arXiv:2509.21828cs.LGcs.MA2025-09被引 1

用偏好学习提升稀疏奖励下的多智能体强化学习效果

Preference-Guided Learning for Sparse-Reward Multi-Agent Reinforcement Learning

  • 通过偏好驱动的值分解网络隐式学习全局与局部奖励信号
  • 在MAMuJoCo和SMACv2上超越现有基线,显著提升性能
  • 可结合大语言模型生成偏好标签,增强奖励学习质量

我们研究在线多智能体强化学习(MARL)中稀疏奖励环境下的问题,其中奖励仅在轨迹结束时反馈,而非每步都提供。这种设置虽真实,但缺乏中间奖励使标准MARL算法难以有效指导策略学习。为此,我们提出一个新框架,将在线逆偏好学习与多智能体在线策略优化统一整合。核心是基于偏好值分解网络构建隐式多智能体奖励学习模型,生成全局与局部奖励信号,并据此构建双优势流,为集中式评估器和分布式执行者提供差异化学习目标。此外,我们展示如何利用大语言模型(LLMs)生成偏好标签以提升学习到的奖励模型质量。在MAMuJoCo和SMACv2等前沿基准上的实验表明,该方法在稀疏奖励在线MARL中表现优于现有基线,验证了其有效性。

原文摘要 · Abstract (English)

We study the problem of online multi-agent reinforcement learning (MARL) in environments with sparse rewards, where reward feedback is not provided at each interaction but only revealed at the end of a trajectory. This setting, though realistic, presents a fundamental challenge: the lack of intermediate rewards hinders standard MARL algorithms from effectively guiding policy learning. To address this issue, we propose a novel framework that integrates online inverse preference learning with multi-agent on-policy optimization into a unified architecture. At its core, our approach introduces an implicit multi-agent reward learning model, built upon a preference-based value-decomposition network, which produces both global and local reward signals. These signals are further used to construct dual advantage streams, enabling differentiated learning targets for the centralized critic and decentralized actors. In addition, we demonstrate how large language models (LLMs) can be leveraged to provide preference labels that enhance the quality of the learned reward model. Empirical evaluations on state-of-the-art benchmarks, including MAMuJoCo and SMACv2, show that our method achieves superior performance compared to existing baselines, highlighting its effectiveness in addressing sparse-reward challenges in online MARL.

多智能体强化学习稀疏奖励偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。