arXiv:2505.21639cs.LGmath.OC2025-05

用逆优化框架统一强化学习与先验信念,提升专家策略学习的稳定性。

Apprenticeship learning with prior beliefs using inverse optimization

  • 将先验信念融入逆强化学习,构建带正则化的极小极大优化模型
  • 正则项有效缓解逆强化学习的病态问题,提升成本函数学习精度
  • 适用于有不完美专家数据的场景,适合机器人控制与智能决策研究者

逆强化学习(IRL)与马尔可夫决策过程(MDP)的逆优化(IO)之间的关系在文献中尚未充分探讨,尽管二者解决的是同一类问题。本文重新审视了基于MDP的逆优化框架、逆强化学习与导师学习(AL)之间的联系。通过在IRL和AL问题中引入对成本函数结构的先验信念,我们发现:从凸分析视角看,传统导师学习形式是本框架的一个松弛版本。值得注意的是,当无正则化项时,导师学习即为该框架的特例。针对次优专家数据情形,我们将导师学习建模为一个带正则化的极小极大问题,其中正则项在克服逆强化学习病态性方面起关键作用,引导寻找合理的成本函数。为求解由此产生的正则化凸-凹极小极大问题,采用随机镜面下降法(SMD),并建立该方法的收敛性边界。数值实验验证了正则化在学习成本向量和模仿策略中的决定性作用。

原文摘要 · Abstract (English)

The relationship between inverse reinforcement learning (IRL) and inverse optimization (IO) for Markov decision processes (MDPs) has been relatively underexplored in the literature, despite addressing the same problem. In this work, we revisit the relationship between the IO framework for MDPs, IRL, and apprenticeship learning (AL). We incorporate prior beliefs on the structure of the cost function into the IRL and AL problems, and demonstrate that the convex-analytic view of the AL formalism emerges as a relaxation of our framework. Notably, the AL formalism is a special case in our framework when the regularization term is absent. Focusing on the suboptimal expert setting, we formulate the AL problem as a regularized min-max problem. The regularizer plays a key role in addressing the ill-posedness of IRL by guiding the search for plausible cost functions. To solve the resulting regularized-convex-concave-min-max problem, we use stochastic mirror descent (SMD) and establish convergence bounds for the proposed method. Numerical experiments highlight the critical role of regularization in learning cost vectors and apprentice policies.

逆强化学习优化先验信念极小极大

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。