提出可有效量化不确定性的逆强化学习新方法,让灵活模型也能做统计推断。
Efficient Inference for Inverse Reinforcement Learning and Dynamic Discrete Choice Models
- 用行为策略当伪奖励,实现奖励的唯一识别
- 构建高效估计器,满足√n一致性与渐近正态性
- 适合需要可靠统计推断的复杂决策建模场景
在许多序列决策问题中,研究者只能观测到行动而无法得知驱动行为的奖励,但仍需评估和比较反事实政策。逆强化学习(IRL)和动态离散选择(DDC)模型通过假设最优性机制将隐含奖励与可观测行动关联。现有灵活的IRL方法虽能表示复杂奖励,但通常无法提供有效的统计推断;经典DDC方法虽支持推断,但仅适用于严格参数化结构。本文提出一种半参数框架,用于最大熵IRL和Gumbel扰动DDC模型中的无偏逆强化学习。核心识别结果是:对数行为策略可作为伪奖励,点识别政策价值差异,并在归一化约束下识别奖励本身。这将奖励相关估计量的推断转化为对行为策略与转移核的光滑函数推断。我们建立了路径可微性,推导出高效影响函数,并构造了自动去偏机器学习估计器,允许灵活地估计干扰项,同时保证√n一致性、渐近正态性和半参数效率。该方法为灵活的IRL和DDC模型提供了计算上可行的有效不确定性量化框架。
原文摘要 · Abstract (English)
In many sequential decision-making problems, researchers observe actions but not the rewards that drive behavior, yet still wish to evaluate and compare counterfactual policies. Inverse reinforcement learning (IRL) and dynamic discrete choice (DDC) models address this setting by positing an optimality model that links latent rewards to observed actions. Existing flexible IRL methods allow rich reward representations but typically do not provide valid inference, whereas classical DDC methods support inference only under restrictive parametric structure. We develop a semiparametric framework for debiased inverse reinforcement learning in maximum-entropy IRL and Gumbel-shock DDC models. Our key identification result is that the log-behavior policy can be treated as a pseudo-reward: it point-identifies policy value differences and, under a normalization constraint, the reward itself. This reduces inference on reward-dependent estimands to inference on smooth functionals of the behavior policy and transition kernel. We establish pathwise differentiability, derive efficient influence functions, and construct automatic debiased machine-learning estimators that permit flexible nuisance estimation while attaining $\sqrt{n}$-consistency, asymptotic normality, and semiparametric efficiency. The result is a computationally tractable framework for valid uncertainty quantification in flexible IRL and DDC models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。