提出奖励兼容性框架,让逆强化学习在复杂场景中更高效可靠。
Reward Compatibility: A Framework for Inverse RL
- 用新框架量化奖励与专家演示的匹配程度,突破传统可行性限制。
- 在多种数据和策略条件下,均给出可计算的算法与样本复杂度分析。
- 适合研究逆强化学习理论或需要高可靠性奖励设计的研究者。
本文从奖励兼容性这一新框架出发,对逆强化学习(IRL)进行原创性理论研究,用于量化奖励与给定专家演示之间的匹配程度。直观上,若某奖励下推导出的专家策略性能越接近该奖励下的最优性能,则该奖励越兼容。这一框架将传统“可行奖励集”中二元兼容性(是/否)扩展为连续度量,使可证明高效的逆强化学习从表格型马尔可夫决策过程推广至大规模问题。我们分析了包括最优与次优专家演示、在线与离线数据收集在内的多种设置,对每种情况都提供了可计算的算法与对应的样本复杂度分析,并揭示了奖励兼容性的深层机制,为更广泛的问题设定奠定基础。
原文摘要 · Abstract (English)
We provide an original theoretical study of Inverse Reinforcement Learning (IRL) through the lens of reward compatibility, a novel framework to quantify the compatibility of a reward with the given expert's demonstrations. Intuitively, a reward is more compatible with the demonstrations the closer the performance of the expert's policy computed with that reward is to the optimal performance for that reward. This generalizes the notion of feasible reward set, the most common framework in the theoretical IRL literature, for which a reward is either compatible or not compatible. The grayscale introduced by the reward compatibility is the key to extend the realm of provably efficient IRL far beyond what is attainable with the feasible reward set: from tabular to large-scale MDPs. We analyze the IRL problem across various settings, including optimal and suboptimal expert's demonstrations and both online and offline data collection. For all of these dimensions, we provide a tractable algorithm and corresponding sample complexity analysis, as well as various insights on reward compatibility and how the framework can pave the way to yet more general problem settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。