arXiv:2411.09891cs.LGcs.AI2024-11NeurIPS被引 20

通过域适应与奖励增强模仿学习,提升动态变化下的策略迁移性能。

Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented Imitation

  • 用奖励修改实现源域到目标域的分布对齐,再通过模仿学习迁移策略。
  • 在多个基准环境上超越纯奖励修改方法和基线模型,性能显著提升。
  • 适合需要跨动态场景迁移策略的研究者,尤其关注强化学习泛化能力。

在源域训练的策略部署到存在动力学差异的目标域时,常因性能下降而失效。已有方法通过修改奖励使源域最优轨迹分布与目标域匹配,但仅保证行为相似性,无法确保实际部署时达到最优。本文提出域适应与奖励增强模仿学习(DARAIL),利用奖励修改进行域适应,并采用生成对抗模仿学习框架(GAIfO)中的奖励增强估计器优化策略。理论上,在温和的动力学偏移假设下,我们给出了方法的误差界,验证了其合理性。实验表明,该方法在多个基准离线动态环境任务中优于纯奖励修改方法及其他基线模型。

原文摘要 · Abstract (English)

Training a policy in a source domain for deployment in the target domain under a dynamics shift can be challenging, often resulting in performance degradation. Previous work tackles this challenge by training on the source domain with modified rewards derived by matching distributions between the source and the target optimal trajectories. However, pure modified rewards only ensure the behavior of the learned policy in the source domain resembles trajectories produced by the target optimal policies, which does not guarantee optimal performance when the learned policy is actually deployed to the target domain. In this work, we propose to utilize imitation learning to transfer the policy learned from the reward modification to the target domain so that the new policy can generate the same trajectories in the target domain. Our approach, Domain Adaptation and Reward Augmented Imitation Learning (DARAIL), utilizes the reward modification for domain adaptation and follows the general framework of generative adversarial imitation learning from observation (GAIfO) by applying a reward augmented estimator for the policy optimization step. Theoretically, we present an error bound for our method under a mild assumption regarding the dynamics shift to justify the motivation of our method. Empirically, our method outperforms the pure modified reward method without imitation learning and also outperforms other baselines in benchmark off-dynamics environments.

强化学习策略迁移域适应模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。