提出新方法让多智能体逆强化学习更稳定可靠
On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning
- 用熵正则化构建唯一均衡,解决奖励歧义问题
- 理论证明样本量与策略性能误差的关系
- 适合研究多智能体协作或博弈的学者参考
多智能体逆强化学习(MAIRL)旨在从专家示范中恢复智能体的奖励函数。本文刻画了马尔可夫博弈中的可行奖励集,识别出所有能解释给定均衡的奖励函数。然而,基于均衡的观测常存在歧义:单一纳什均衡可能对应多种奖励结构,从而改变多智能体系统的本质。为此,我们引入熵正则化马尔可夫博弈,在保持战略激励的同时获得唯一均衡。针对该设定,我们提供了样本复杂性分析,详细说明误差如何影响学习到的策略性能。本工作建立了MAIRL的理论基础并提供实用洞见。
原文摘要 · Abstract (English)
Multi-agent Inverse Reinforcement Learning (MAIRL) aims to recover agent reward functions from expert demonstrations. We characterize the feasible reward set in Markov games, identifying all reward functions that rationalize a given equilibrium. However, equilibrium-based observations are often ambiguous: a single Nash equilibrium can correspond to many reward structures, potentially changing the game's nature in multi-agent systems. We address this by introducing entropy-regularized Markov games, which yield a unique equilibrium while preserving strategic incentives. For this setting, we provide a sample complexity analysis detailing how errors affect learned policy performance. Our work establishes theoretical foundations and practical insights for MAIRL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。