arXiv:2605.27834cs.LGstat.ML2026-05

通过联合求解双环境贝尔曼方程,提升逆强化学习奖励迁移效果。

Reward Transfer from Inverse Reinforcement Learning: A Coupled Minimax Approach

  • 构建源环境与目标环境的耦合贝尔曼方程系统,统一求解目标软Q函数。
  • 理论证明该方法可消除源环境贝尔曼残差的一阶影响,误差更小。
  • 适用于从受控环境迁移奖励到新场景的强化学习任务,如医疗决策模拟。

研究如何将专家演示中通过逆强化学习获得的奖励,从一个环境迁移到另一个不同环境的强化学习任务中。当演示在受控环境中收集时,此问题自然出现。本文将问题建模为源环境与目标环境间的联合贝尔曼方程系统,并提出目标软Q函数的极小极大估计器。与先估计源奖励再用于目标控制的顺序方法不同,耦合方法联合求解源与目标方程。理论表明,相较于顺序方法,耦合方法消除了源贝尔曼残差的一阶影响。本文分析了两种方法的局部行为,推导出有限样本下软Q函数的误差界,并证明了所得软控制策略的后悔上界。基于脓毒症模拟器的实验验证了理论比较结果。

原文摘要 · Abstract (English)

We study the transfer of rewards learned using inverse reinforcement learning from expert demonstrations in one environment to reinforcement learning in a new, different environment. This arises naturally when demonstrations are collected in a controlled environment. We formulate the problem as a joint system of Bellman equations across the source and target environments and develop minimax estimators for the target soft-$q$-function. Whereas a sequential solution approach first estimates the source reward and then plugs it into the target control problem, a coupled approach solves the source and target system of equations jointly. We show that, in contrast to the sequential approach, the coupled approach removes the first-order influence of source Bellman residual error. We characterize the local behavior of each approach, develop finite-sample soft-$q$-function error bounds, and prove regret guarantees for the resulting soft-control policy. An empirical investigation using a sepsis simulator validates the theoretical comparison.

奖励迁移逆强化学习强化学习极小极大

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。