arXiv:2503.17865stat.MLcs.LG2025-03被引 1

首次实现神经网络奖励下的全局最优逆强化学习,且有非渐近收敛保证。

Understanding Inverse Reinforcement Learning under Overparameterization: Non-Asymptotic Analysis and Global Optimality

  • 设计双时间尺度单循环算法,突破传统嵌套结构效率瓶颈。
  • 在过参数化条件下证明算法收敛至全局最优奖励与策略。
  • 适用于高维复杂任务,为神经网络逆强化学习提供理论保障。

逆强化学习(IRL)的目标是从专家示范中识别出底层奖励函数及其对应最优策略。尽管大多数IRL算法的理论保证依赖于线性奖励结构,本文将理论理解拓展至由神经网络参数化的奖励场景。传统IRL算法通常采用嵌套结构,导致计算效率低下,尤其在高维设置下更为明显。为此,本文提出首个基于神经网络奖励的双时间尺度单循环IRL算法,并在过参数化条件下提供了非渐近收敛分析。尽管针对线性奖励的先前最优性结果不再适用,我们证明该算法可在特定神经网络结构下识别出全局最优的奖励函数与策略。这是首个在神经网络设定下具有非渐近收敛保证并能严格实现全局最优性的IRL算法。

原文摘要 · Abstract (English)

The goal of the Inverse reinforcement learning (IRL) task is to identify the underlying reward function and the corresponding optimal policy from a set of expert demonstrations. While most IRL algorithms' theoretical guarantees rely on a linear reward structure, we aim to extend the theoretical understanding of IRL to scenarios where the reward function is parameterized by neural networks. Meanwhile, conventional IRL algorithms usually adopt a nested structure, leading to computational inefficiency, especially in high-dimensional settings. To address this problem, we propose the first two-timescale single-loop IRL algorithm under neural network parameterized reward and provide a non-asymptotic convergence analysis under overparameterization. Although prior optimality results for linear rewards do not apply, we show that our algorithm can identify the globally optimal reward and policy under certain neural network structures. This is the first IRL algorithm with a non-asymptotic convergence guarantee that provably achieves global optimality in neural network settings.

逆强化学习神经网络全局最优非渐近分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。