解决强化学习逆问题中奖励函数不唯一难题,给出可量化收敛的统计框架。
Statistical analysis of Inverse Entropy-regularized Reinforcement Learning
- 用熵正则化+最小二乘重构,确保恢复奖励唯一
- 理论证明收敛速度达到极小极大最优,依赖样本量与模型复杂度
- 适合研究逆强化学习理论或需稳定奖励推断的研究者
逆强化学习旨在从专家轨迹中推断解释其行为的奖励函数。经典逆强化学习长期面临奖励函数不唯一的问题:多个不同奖励可能诱导相同最优策略,导致逆问题病态。本文提出一种基于熵正则化的逆强化学习统计框架,通过结合熵正则化与软贝尔曼残差的最小二乘重构,获得唯一且定义良好的最小二乘奖励。将专家示范建模为马尔可夫链,其不变分布由未知专家策略 $π^ ext{⋆}$ 定义,通过在动作空间条件分布类上使用惩罚最大似然估计策略。建立了估计策略与专家策略之间超出Kullback-Leibler散度的高概率界,统计复杂度由策略类的覆盖数刻画。结果揭示了平滑(熵正则化)、模型复杂度与样本量之间的权衡关系,并给出了最小二乘奖励函数的非渐近极小极大最优收敛率。本分析连接了行为克隆、逆强化学习与现代统计学习理论。
原文摘要 · Abstract (English)
Inverse reinforcement learning aims to infer the reward function that explains expert behavior observed through trajectories of state--action pairs. A long-standing difficulty in classical IRL is the non-uniqueness of the recovered reward: many reward functions can induce the same optimal policy, rendering the inverse problem ill-posed. In this paper, we develop a statistical framework for Inverse Entropy-regularized Reinforcement Learning that resolves this ambiguity by combining entropy regularization with a least-squares reconstruction of the reward from the soft Bellman residual. This combination yields a unique and well-defined so-called least-squares reward consistent with the expert policy. We model the expert demonstrations as a Markov chain with the invariant distribution defined by an unknown expert policy $π^\star$ and estimate the policy by a penalized maximum-likelihood procedure over a class of conditional distributions on the action space. We establish high-probability bounds for the excess Kullback--Leibler divergence between the estimated policy and the expert policy, accounting for statistical complexity through covering numbers of the policy class. These results lead to non-asymptotic minimax optimal convergence rates for the least-squares reward function, revealing the interplay between smoothing (entropy regularization), model complexity, and sample size. Our analysis bridges the gap between behavior cloning, inverse reinforcement learning, and modern statistical learning theory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。