提出逆强化学习的快速收敛理论,证明样本量越大效果越好
Fast Rates for Inverse Reinforcement Learning
- 用极值优化方法统一最大似然估计,理论更严密
- 专家轨迹数每增加一倍,误差下降至原来的1/4
- 无需覆盖所有状态也能保证结果可靠,适合实际应用
我们在有限时域马尔可夫决策过程(MDP)中,针对具有Borel状态与动作空间的熵正则化极小极大逆强化学习(Min-Max-IRL),建立了新的结构与统计结果。证明了在总体层面,最大似然估计(MLE)与Min-Max-IRL等价;在确定性动态下,经验层面亦等价。对于线性奖励类,利用最小极大损失的伪自洽性,证明了轨迹级KL散度超额损失和参数误差在海森矩阵范数下的衰减速率达到快率$O(n^{-1})$,其中 $n$ 为专家轨迹数量。在设定正确且确定性的局部极小极大下界,匹配参数误差率,仅差对数因子。我们的保证在模型误设下依然成立,且无需均匀状态覆盖假设。进一步将奖励可识别性结果推广至一般Borel空间,并与基于MLE的保证进行了比较。
原文摘要 · Abstract (English)
We establish novel structural and statistical results for entropy-regularized min-max inverse reinforcement learning (Min-Max-IRL) in finite-horizon MDPs with Borel state and action spaces. We show that maximum likelihood estimation (MLE) and Min-Max-IRL are equivalent at the population level, and at the empirical level under deterministic dynamics. For linear reward classes, we leverage pseudo-self-concordance of the Min-Max-IRL loss to prove that both the excess trajectory-level KL divergence and the squared parameter error in the Hessian norm decay at the fast rate $O(n^{-1})$, where $n$ is the number of expert trajectories. A local minimax lower bound matches the parameter-error rate up to logarithmic factors in the well-specified deterministic setting. Our guarantees apply under misspecification and require no uniform state-coverage assumption. We further extend reward-identifiability results to general Borel spaces and compare our results with MLE-based guarantees.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。