arXiv:2509.12010cs.LGcs.AI2025-09被引 1

通过闭式奖励中心点实现专家行为的泛化,提升迁移可靠性。

Generalizing Behavior via Inverse Reinforcement Learning with Closed-Form Reward Centroids

  • 基于可行奖励集的平均策略,用中心点奖励直接规划
  • 推导出奖励中心点的闭式表达式,支持高效计算
  • 仅需离线演示数据即可估计,适合新环境迁移场景

本文研究如何将专家代理的行为(由示范提供)泛化到新环境和额外约束下。逆强化学习(IRL)通过恢复专家的潜在奖励函数来实现这一目标,若在新环境中使用该奖励进行规划,可重现期望行为。然而,IRL本质上是病态问题:多个奖励函数可解释相同行为,形成所谓的可行集。这些奖励在新环境下可能诱导不同策略,缺乏额外信息时需选择部署哪个策略。本文提出一种新的、有原则的决策准则,从可行集的某个有界子集中选择“平均”策略。令人惊讶的是,该策略可通过使用该子集奖励中心点进行规划获得,我们推导出了该中心点的闭式表达式。随后,我们提出一种仅使用离线专家示范数据的、可证明高效的算法来估计该中心点。最后,数值模拟展示了专家行为与本文方法生成行为之间的关系。

原文摘要 · Abstract (English)

We study the problem of generalizing an expert agent's behavior, provided through demonstrations, to new environments and/or additional constraints. Inverse Reinforcement Learning (IRL) offers a promising solution by seeking to recover the expert's underlying reward function, which, if used for planning in the new settings, would reproduce the desired behavior. However, IRL is inherently ill-posed: multiple reward functions, forming the so-called feasible set, can explain the same observed behavior. Since these rewards may induce different policies in the new setting, in the absence of additional information, a decision criterion is needed to select which policy to deploy. In this paper, we propose a novel, principled criterion that selects the "average" policy among those induced by the rewards in a certain bounded subset of the feasible set. Remarkably, we show that this policy can be obtained by planning with the reward centroid of that subset, for which we derive a closed-form expression. We then present a provably efficient algorithm for estimating this centroid using an offline dataset of expert demonstrations only. Finally, we conduct numerical simulations that illustrate the relationship between the expert's behavior and the behavior produced by our method.

逆强化学习行为泛化奖励学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。