从延迟干扰的专家轨迹中恢复奖励特征,提升模仿学习效果
Inverse Delayed Reinforcement Learning
- 用对抗性离线训练从延迟观测中提取专家特征
- 在MuJoCo环境下多种延迟设置下表现优于直接使用延迟观测
- 理论证明增强延迟观测比原始延迟观测更利于政策恢复
逆强化学习(IRL)在多种模仿任务中表现出色。本文提出一种新型IRL框架,旨在从受延迟干扰影响的专家轨迹中提取奖励特征。不同于依赖直接观测的方法,本方法采用高效的离线策略对抗训练框架,从增强的延迟观测中推导专家特征并恢复最优策略。在多种延迟设置下的MuJoCo环境中的实证评估验证了该方法的有效性。此外,我们提供了理论分析,表明利用增强延迟观测恢复专家策略的表现优于使用原始延迟观测。
原文摘要 · Abstract (English)
Inverse Reinforcement Learning (IRL) has demonstrated effectiveness in a variety of imitation tasks. In this paper, we introduce an IRL framework designed to extract rewarding features from expert trajectories affected by delayed disturbances. Instead of relying on direct observations, our approach employs an efficient off-policy adversarial training framework to derive expert features and recover optimal policies from augmented delayed observations. Empirical evaluations in the MuJoCo environment under diverse delay settings validate the effectiveness of our method. Furthermore, we provide a theoretical analysis showing that recovering expert policies from augmented delayed observations outperforms using direct delayed observations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。