arXiv:2410.03847cs.LGcs.AI2024-10被引 4

将环境动态信息融入奖励设计,提升强化学习逆向求解在随机环境中的表现。

Enhancing Inverse Reinforcement Learning through Encoding Dynamic Information in Reward Shaping

  • 通过引入环境转移模型增强奖励函数,改进传统逆强化学习框架。
  • 在MuJoCo随机环境中性能显著优于基线,样本效率大幅提升。
  • 理论证明了奖励误差与策略性能差距的边界,适用于复杂随机场景。

本文针对对抗式逆强化学习(AIRL)在随机环境中的理论失效与性能下降问题,提出一种新方法,将环境动态信息编码至奖励函数中,并为随机环境下诱导出的最优策略提供理论保障。所提出的模型增强型逆强化学习框架(Model-Enhanced AIRL)直接将转移模型估计融入奖励设计。我们还对方法的奖励误差界与性能差异界进行了全面理论分析。在MuJoCo基准测试中,该方法在随机环境中表现更优,在确定性环境中也具有竞争力,且样本效率显著提升。

原文摘要 · Abstract (English)

In this paper, we aim to tackle the limitation of the Adversarial Inverse Reinforcement Learning (AIRL) method in stochastic environments where theoretical results cannot hold and performance is degraded. To address this issue, we propose a novel method which infuses the dynamics information into the reward shaping with the theoretical guarantee for the induced optimal policy in the stochastic environments. Incorporating our novel model-enhanced rewards, we present a novel Model-Enhanced AIRL framework, which integrates transition model estimation directly into reward shaping. Furthermore, we provide a comprehensive theoretical analysis of the reward error bound and performance difference bound for our method. The experimental results in MuJoCo benchmarks show that our method can achieve superior performance in stochastic environments and competitive performance in deterministic environments, with significant improvement in sample efficiency, compared to existing baselines.

逆强化学习奖励设计随机环境模型增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。