基于人类偏好的多智能体强化学习新框架,提升训练稳定性与效率。
O-MAPL: Offline Multi-agent Preference Learning
- 端到端融合偏好学习与价值分解,直接优化联合策略。
- 在SMAC和MAMuJoCo上超越现有方法,性能提升显著。
- 适合需要高效协作的多智能体系统研究者使用。
从示范中推断奖励函数是强化学习中的关键挑战,尤其在多智能体强化学习(MARL)中,庞大的联合状态-动作空间和复杂的智能体间交互使该任务更加困难。尽管已有单智能体研究探索了从人类偏好中恢复奖励函数与策略,但针对MARL的研究仍有限。现有方法通常分阶段进行监督式奖励学习与MARL算法,导致训练不稳定。本文提出一种新型端到端偏好驱动的协作式MARL学习框架,利用奖励函数与软Q函数之间的内在关联,采用精心设计的多智能体价值分解策略,提升训练效率。在SMAC与MAMuJoCo基准上的大量实验表明,该算法在多种任务中均优于现有方法。
原文摘要 · Abstract (English)
Inferring reward functions from demonstrations is a key challenge in reinforcement learning (RL), particularly in multi-agent RL (MARL), where large joint state-action spaces and complex inter-agent interactions complicate the task. While prior single-agent studies have explored recovering reward functions and policies from human preferences, similar work in MARL is limited. Existing methods often involve separate stages of supervised reward learning and MARL algorithms, leading to unstable training. In this work, we introduce a novel end-to-end preference-based learning framework for cooperative MARL, leveraging the underlying connection between reward functions and soft Q-functions. Our approach uses a carefully-designed multi-agent value decomposition strategy to improve training efficiency. Extensive experiments on SMAC and MAMuJoCo benchmarks show that our algorithm outperforms existing methods across various tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。