提出MODULE算法,解决观察学习中的样本效率与训练不稳问题。
On Generalization and Distributional Update for Mimicking Observations with Adequate Exploration
- 结合分布强化学习与软演员-评论家,提升策略优化稳定性。
- 在MuJoCo环境上显著优于现有观察学习方法。
- 适合需要高效稳定模仿专家行为的研究者使用。
观察学习(LfO)无需访问专家动作即可复现其行为,在真实场景中比示范学习更实用。然而,直接采用在线策略训练会加剧样本低效,传统离线策略训练又放大了训练不稳定性。本文通过分析奖励函数与策略的泛化能力,构建理论基础,并改进生成对抗式观察模仿(GAIfO)中的策略优化方法,提出模块化更新学习算法(MODULE)。该算法融合软演员-评论家(SAC)的高样本效率与训练鲁棒性,以及分布强化学习的训练稳定性优势。在MuJoCo环境中的大量实验表明,MODULE显著优于现有观察学习方法。
原文摘要 · Abstract (English)
Learning from observations (LfO) replicates expert behavior without needing access to the expert's actions, making it more practical than learning from demonstrations (LfD) in many real-world scenarios. However, directly applying the on-policy training scheme in LfO worsens the sample inefficiency problem, while employing the traditional off-policy training scheme in LfO magnifies the instability issue. This paper seeks to develop an efficient and stable solution for the LfO problem. Specifically, we begin by exploring the generalization capabilities of both the reward function and policy in LfO, which provides a theoretical foundation for computation. Building on this, we modify the policy optimization method in generative adversarial imitation from observation (GAIfO) with distributional soft actor-critic (DSAC), and propose the Mimicking Observations through Distributional Update Learning with adequate Exploration (MODULE) algorithm to solve the LfO problem. MODULE incorporates the advantages of (1) high sample efficiency and training robustness enhancement in soft actor-critic (SAC), and (2) training stability in distributional reinforcement learning (RL). Extensive experiments in MuJoCo environments showcase the superior performance of MODULE over current LfO methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。