揭示强化学习逆问题中的奖励函数不确定性与模型偏差根源
Partial Identifiability and Misspecification in Inverse Reinforcement Learning
- 系统分析逆强化学习中奖励函数的不完全可识别性
- 给出常见行为模型在偏差下的误判边界条件
- 提供可推广的数学框架,适用于新模型验证
逆强化学习(IRL)旨在从策略π中推断奖励函数R,但该问题具有根本性挑战。首先,通常存在多个与同一策略兼容的奖励函数,导致奖励函数仅部分可识别,蕴含固有模糊性。其次,推断R需依赖行为模型来描述π与R的关系,而真实人类偏好与行为间关系极为复杂,现有模型难以完全刻画,因此行为模型在实践中普遍存在错设。本文对逆强化学习中的部分可识别性与模型错设进行了全面数学分析,完整刻画并量化了当前主流行为模型下奖励函数的模糊程度;给出了标准行为模型在何种条件下会导致对奖励函数的错误推断的充要条件。此外,提出一个统一的分析框架及若干形式化工具,可便捷推导新IRL模型的可识别性与错设鲁棒性,或用于分析其他奖励学习算法。
原文摘要 · Abstract (English)
The aim of Inverse Reinforcement Learning (IRL) is to infer a reward function $R$ from a policy $π$. This problem is difficult, for several reasons. First of all, there are typically multiple reward functions which are compatible with a given policy; this means that the reward function is only *partially identifiable*, and that IRL contains a certain fundamental degree of ambiguity. Secondly, in order to infer $R$ from $π$, an IRL algorithm must have a *behavioural model* of how $π$ relates to $R$. However, the true relationship between human preferences and human behaviour is very complex, and practically impossible to fully capture with a simple model. This means that the behavioural model in practice will be *misspecified*, which raises the worry that it might lead to unsound inferences if applied to real-world data. In this paper, we provide a comprehensive mathematical analysis of partial identifiability and misspecification in IRL. Specifically, we fully characterise and quantify the ambiguity of the reward function for all of the behavioural models that are most common in the current IRL literature. We also provide necessary and sufficient conditions that describe precisely how the observed demonstrator policy may differ from each of the standard behavioural models before that model leads to faulty inferences about the reward function $R$. In addition to this, we introduce a cohesive framework for reasoning about partial identifiability and misspecification in IRL, together with several formal tools that can be used to easily derive the partial identifiability and misspecification robustness of new IRL models, or analyse other kinds of reward learning algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。