arXiv:2409.00717cs.LGcs.AI2024-09被引 1

提出偏好驱动的多智能体强化学习新框架,解决稀疏反馈难题。

Preference-Based Multi-Agent Reinforcement Learning: Data Coverage and Algorithmic Techniques

  • 基于偏好数据构建博弈均衡求解方法,强调单策略覆盖不足。
  • 理论证明需单边数据覆盖才能有效逼近纳什均衡。
  • 设计时序正则与分布惩罚机制,提升训练稳定性和性能。

我们开创性地研究了基于偏好的多智能体强化学习(PbMARL),探索其理论基础与实证验证。将任务定义为:从仅含偏好的离线数据集中识别一般和博弈中的纳什均衡,该问题因反馈信号稀疏而极具挑战。理论分析揭示了有效PbMARL中纳什均衡的上界复杂度,表明单一策略覆盖不足,凸显单边数据覆盖的重要性。这些理论发现通过全面实验得到验证。为提升实际性能,我们进一步提出两种算法技术:(1) 在时间轴上引入均方误差(MSE)正则化,实现更均匀的奖励分布,改善奖励学习效果;(2) 基于数据集分布添加额外惩罚项,引入悲观性,增强训练过程中的稳定性和有效性。研究结果强调了PbMARL所需多维度方法,为高效偏好驱动的多智能体系统奠定基础。

原文摘要 · Abstract (English)

We initiate the study of Preference-Based Multi-Agent Reinforcement Learning (PbMARL), exploring both theoretical foundations and empirical validations. We define the task as identifying the Nash equilibrium from a preference-only offline dataset in general-sum games, a problem marked by the challenge of sparse feedback signals. Our theory establishes the upper complexity bounds for Nash Equilibrium in effective PbMARL, demonstrating that single-policy coverage is inadequate and highlighting the importance of unilateral dataset coverage. These theoretical insights are verified through comprehensive experiments. To enhance the practical performance, we further introduce two algorithmic techniques. (1) We propose a Mean Squared Error (MSE) regularization along the time axis to achieve a more uniform reward distribution and improve reward learning outcomes. (2) We propose an additional penalty based on the distribution of the dataset to incorporate pessimism, improving stability and effectiveness during training. Our findings underscore the multifaceted approach required for PbMARL, paving the way for effective preference-based multi-agent systems.

多智能体偏好学习纳什均衡强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。