arXiv:2502.14131cs.LGcs.AI2025-02被引 1

无需假设奖励线性,用经验风险最小化高效求解离线决策模型。

An Empirical Risk Minimization Approach for Offline Inverse RL and Dynamic Discrete Choice Model

  • 基于经验风险最小化构建逆强化学习框架,避开状态转移概率估计
  • 在合成数据上显著优于基准方法,且支持神经网络非参数建模
  • 适用于高维、无限状态空间,理论保证快速全局收敛

本文研究动态离散选择(DDC)模型的估计问题,即从离线行为数据中恢复决定代理行为的奖励函数或$Q^*$函数,也称为机器学习中的离线最大熵正则逆强化学习(offline MaxEnt-IRL)。提出一种全局收敛的基于梯度的方法,无需对奖励函数作线性参数化假设。新方法的核心在于引入基于经验风险最小化(ERM)的逆强化学习/动态离散选择框架,避免了贝尔曼方程中显式状态转移概率的估计。此外,该方法兼容神经网络等非参数估计技术,具备扩展至高维、无限状态空间的潜力。关键理论发现是贝尔曼残差满足Polyak-Lojasiewicz(PL)条件——虽弱于强凸性,但足以保证快速全局收敛。通过一系列合成实验验证,所提方法持续优于基准方法和最先进算法。

原文摘要 · Abstract (English)

We study the problem of estimating Dynamic Discrete Choice (DDC) models, also known as offline Maximum Entropy-Regularized Inverse Reinforcement Learning (offline MaxEnt-IRL) in machine learning. The objective is to recover reward or $Q^*$ functions that govern agent behavior from offline behavior data. In this paper, we propose a globally convergent gradient-based method for solving these problems without the restrictive assumption of linearly parameterized rewards. The novelty of our approach lies in introducing the Empirical Risk Minimization (ERM) based IRL/DDC framework, which circumvents the need for explicit state transition probability estimation in the Bellman equation. Furthermore, our method is compatible with non-parametric estimation techniques such as neural networks. Therefore, the proposed method has the potential to be scaled to high-dimensional, infinite state spaces. A key theoretical insight underlying our approach is that the Bellman residual satisfies the Polyak-Lojasiewicz (PL) condition -- a property that, while weaker than strong convexity, is sufficient to ensure fast global convergence guarantees. Through a series of synthetic experiments, we demonstrate that our approach consistently outperforms benchmark methods and state-of-the-art alternatives.

逆强化学习动态决策经验风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。