通过联合预训练策略与奖励,显著提升对抗性模仿学习的效率与效果。
Provably Efficient Policy-Reward Co-Pretraining for Adversarial Imitation Learning

- 提出策略与奖励联合预训练新方法,解决仅预训练策略的局限。
- 理论证明新方法可缩小模仿差距,优于传统对抗模仿学习。
- 适合研究模仿学习理论或追求高效强化学习算法的研究者。
对抗性模仿学习(AIL)相比行为克隆(BC)能实现更高质量的模仿,但需要大量在线环境交互。近期工作尝试用预训练的BC策略初始化AIL算法以缓解此问题,但缺乏对预训练在AIL中作用的严格理论理解。本文系统分析了仅使用策略预训练的AIL,发现奖励误差是导致性能不佳的主要原因,揭示了此前被忽视的关键缺口:缺乏奖励预训练。基于此,我们提出一种基于奖励塑造分析的原理性联合预训练方法。分析揭示了专家策略与塑造奖励之间的根本联系,自然导出CoPT-AIL,该方法通过一次BC过程联合预训练策略与奖励。我们证明,CoPT-AIL相较于标准AIL实现了更优的模仿差距界,首次为预训练在AIL中的优势提供了理论保证。实验结果验证了CoPT-AIL在性能上优于现有AIL方法。
原文摘要 · Abstract (English)
Adversarial imitation learning (AIL) achieves high-quality imitation compared to behavioral cloning (BC), but demands substantial online environment interaction. Recent empirical work has explored initializing AIL algorithms with BC pretrained policies to address this limitation, yet a rigorous theoretical understanding of pretraining's role in AIL remains elusive. This paper provides a systematic theoretical analysis and introduces principled pretraining algorithms for accelerating AIL. We begin by analyzing AIL with policy pretraining alone, identifying reward error as the dominant source of suboptimality. This reveals a critical and previously overlooked gap: the absence of reward pretraining. Motivated by this finding, we develop a principled policy-reward co-pretraining approach grounded in a reward shaping analysis. Our analysis uncovers a fundamental connection between expert policies and shaping rewards, which naturally gives rise to CoPT-AIL, an approach that jointly pretrains both policy and reward through a single BC procedure. We prove that CoPT-AIL achieves an improved imitation gap bound over standard AIL, establishing the first theoretical guarantee for the benefits of pretraining in AIL. Experimental results confirm CoPT-AIL's superior performance over existing AIL methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。