arXiv:2409.08012cs.LGcs.AI2024-09被引 2

通过因果不变性正则化,提升逆强化学习的泛化能力。

Learning Causally Invariant Reward Functions from Diverse Demonstrations

  • 基于因果不变性原则设计正则化方法,减少数据中的虚假相关。
  • 在分布偏移下训练策略时,性能优于传统方法。
  • 适合需要跨环境迁移的强化学习任务。

逆强化学习旨在根据专家示范数据集恢复马尔可夫决策过程的奖励函数。然而,示范数据常因稀缺性和来源异质性,导致学习到的奖励函数吸收数据中的虚假相关。这使得在环境动态分布偏移下,基于该奖励函数训练的策略表现出行为过拟合。本文提出一种基于因果不变性原则的新正则化方法,应用于精确与近似两种逆强化学习形式,在迁移设置中验证了所恢复奖励函数能带来更优的策略性能。

原文摘要 · Abstract (English)

Inverse reinforcement learning methods aim to retrieve the reward function of a Markov decision process based on a dataset of expert demonstrations. The commonplace scarcity and heterogeneous sources of such demonstrations can lead to the absorption of spurious correlations in the data by the learned reward function. Consequently, this adaptation often exhibits behavioural overfitting to the expert data set when a policy is trained on the obtained reward function under distribution shift of the environment dynamics. In this work, we explore a novel regularization approach for inverse reinforcement learning methods based on the causal invariance principle with the goal of improved reward function generalization. By applying this regularization to both exact and approximate formulations of the learning task, we demonstrate superior policy performance when trained using the recovered reward functions in a transfer setting

逆强化学习因果不变性泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。