arXiv:2509.23135cs.LGcs.AI2025-09NeurIPS

提出稳定高效的逆强化学习算法,提升专家行为模仿效果。

Trust Region Reward Optimization and Proximal Inverse Reward Optimization Algorithm

  • 基于信任区域优化框架,保证奖励与策略联合学习的稳定性
  • 在多个仿真和真实场景中实现高样本效率的策略模仿
  • 理论对齐TRPO,适合需要可靠训练的逆强化学习任务

逆强化学习(IRL)旨在通过专家示范学习解释其行为的奖励函数。现有方法多采用对抗式(极小极大)框架,交替优化奖励与策略,常导致训练不稳定。近期非对抗性方法通过能量模型联合学习奖励与策略,提升了稳定性,但缺乏形式化保证。本文首次统一视角,揭示经典非对抗方法实质上显式或隐式最大化专家行为的似然,等价于最小化期望回报差距。由此提出信任区域奖励优化(TRRO)框架,通过极小化-极大化过程保证该似然的单调提升。进一步实例化为近端逆向奖励优化(PIRO),一个实用且稳定的IRL算法。理论上,TRRO为逆强化学习提供了类似正向强化学习中TRPO的稳定性保障。实验表明,PIRO在MuJoCo、Gym-Robotics基准及真实动物行为建模任务中,均达到或超越当前最优水平,具备优异的奖励恢复能力与高样本效率。

原文摘要 · Abstract (English)

Inverse Reinforcement Learning (IRL) learns a reward function to explain expert demonstrations. Modern IRL methods often use the adversarial (minimax) formulation that alternates between reward and policy optimization, which often lead to unstable training. Recent non-adversarial IRL approaches improve stability by jointly learning reward and policy via energy-based formulations but lack formal guarantees. This work bridges this gap. We first present a unified view showing canonical non-adversarial methods explicitly or implicitly maximize the likelihood of expert behavior, which is equivalent to minimizing the expected return gap. This insight leads to our main contribution: Trust Region Reward Optimization (TRRO), a framework that guarantees monotonic improvement in this likelihood via a Minorization-Maximization process. We instantiate TRRO into Proximal Inverse Reward Optimization (PIRO), a practical and stable IRL algorithm. Theoretically, TRRO provides the IRL counterpart to the stability guarantees of Trust Region Policy Optimization (TRPO) in forward RL. Empirically, PIRO matches or surpasses state-of-the-art baselines in reward recovery, policy imitation with high sample efficiency on MuJoCo and Gym-Robotics benchmarks and a real-world animal behavior modeling task.

逆强化学习信任区域稳定训练策略模仿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。