利用多智能体对称性提升逆强化学习样本效率
Symmetry-Guided Multi-Agent Inverse Reinforcement Learning
- 基于多智能体系统内在对称性设计新框架
- 在少量专家示范下仍能准确恢复奖励函数
- 适合需要高效训练的多机器人实际场景
在机器人系统中,强化学习性能依赖于预设奖励函数的合理性。然而,人工设计的奖励函数常因不准确导致策略失败。逆强化学习(IRL)通过从专家示范中推断隐含奖励函数来解决此问题。但现有方法严重依赖大量专家示范才能准确恢复奖励函数。在机器人应用中,尤其是多机器人系统中,收集专家示范成本高昂,严重制约了IRL的实际部署。因此,提高样本效率已成为多智能体逆强化学习(MIRL)的关键挑战。受多智能体系统固有对称性的启发,本文理论证明:利用对称性可更准确地恢复奖励函数。基于此,我们提出一个通用框架,将对称性整合进现有的多智能体对抗式IRL算法,显著提升样本效率。多个复杂任务的实验结果验证了该框架的有效性。进一步在物理多机器人系统中的验证表明本方法具备实际可行性。
原文摘要 · Abstract (English)
In robotic systems, the performance of reinforcement learning depends on the rationality of predefined reward functions. However, manually designed reward functions often lead to policy failures due to inaccuracies. Inverse Reinforcement Learning (IRL) addresses this problem by inferring implicit reward functions from expert demonstrations. Nevertheless, existing methods rely heavily on large amounts of expert demonstrations to accurately recover the reward function. The high cost of collecting expert demonstrations in robotic applications, particularly in multi-robot systems, severely hinders the practical deployment of IRL. Consequently, improving sample efficiency has emerged as a critical challenge in multi-agent inverse reinforcement learning (MIRL). Inspired by the symmetry inherent in multi-agent systems, this work theoretically demonstrates that leveraging symmetry enables the recovery of more accurate reward functions. Building upon this insight, we propose a universal framework that integrates symmetry into existing multi-agent adversarial IRL algorithms, thereby significantly enhancing sample efficiency. Experimental results from multiple challenging tasks have demonstrated the effectiveness of this framework. Further validation in physical multi-robot systems has shown the practicality of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。