arXiv:2605.30843cs.LGecon.EM2026-05

从专家数据反推其优化的奖励函数,打通经济学与机器学习的桥梁。

A Lecture Note on Offline RL and IRL, Part II: Foundations of Inverse Reinforcement Learning and Dynamic Discrete Choice Models

  • 统一了经济计量学与机器学习中的逆强化学习模型,证明两者本质相同。
  • 揭示了经典方法在维度、转移核估计等方面的局限性,提出可识别性边界。
  • 结合现代方法如对抗式逆强化学习,给出可训练的梯度估计器,适合实操研究者。

在正向强化学习中,奖励函数已知,目标是寻找最优策略或价值函数;而逆强化学习则相反:给定由专家生成的离线数据,能否恢复其优化的奖励函数?这正是逆强化学习的核心问题。令人惊讶的是,结构计量经济学家研究动态离散选择(DDC)与机器学习领域研究熵正则化逆强化学习的两个群体,实际上在使用相同的概率模型,仅名称不同。本文首先证明两者的等价性。随后,系统梳理了Magnac和Thesmar的经典可识别性结果及其衍生的计算范式:Rust的嵌套不动点算法、Hotz与Miller的条件选择概率法,以及Adusumilli与Eckardt提出的两种时序差分方法——线性半梯度TD与近似值迭代。这些方法各有缺陷:维度瓶颈、转移核估计困难、‘致命三重困境’或投影不动点偏差。接着,本文深入分析现代机器学习/逆强化学习路径:对抗式逆强化学习、占据匹配、IQ-Learn及离线机器学习逆强化学习,明确各方法的真实目标及其可识别范围。最后,引入Kang等人提出的经验风险最小化框架,提供一种基于梯度的离线逆强化学习/动态离散选择估计方法。

原文摘要 · Abstract (English)

In the forward reinforcement-learning problem, the reward is fixed and known; the learner is asked to find a good policy or value function. Here we turn the question around. Given offline data generated by an expert, can we recover the reward the expert was optimizing? This is the inverse reinforcement learning problem, and remarkably, two communities, structural econometricians studying dynamic discrete choice (DDC) and machine learners studying entropy-regularized IRL, have been working on exactly the same probabilistic model under different names. We begin by proving their equivalence. We then develop the classical identification result of Magnac and Thesmar and the classical computational paradigms that grew out of it: Rust's nested fixed-point algorithm, the conditional-choice-probability approach of Hotz and Miller, and the two temporal-difference approaches of Adusumilli and Eckardt: linear semi-gradient TD and approximate value iteration. Each route has its limits: dimensionality, transition-kernel estimation, the deadly triad, or projected fixed-point bias. We then walk through the modern ML/IRL strand: adversarial IRL, occupancy matching, IQ-Learn, and offline ML-IRL, deriving each method's actual objective and stating precisely what it does and does not identify. We close with the empirical-risk-minimization framework of Kang et al., which yields a gradient-based estimator for offline IRL/DDC.

逆强化学习动态离散选择离线强化学习经济学+机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。