通过离线到在线的自适应学习,缓解机器人控制中的分布偏移问题。
How to Mitigate the Distribution Shift Problem in Robotics Control: A Robust and Adaptive Approach Based on Offline to Online Imitation Learning

- 用判别器筛选补充演示数据,扩大策略的状态-动作覆盖范围。
- 在线阶段检测分布偏移,并基于自监督模仿学习实时更新策略。
- 在MuJoCo环境中显著提升对分布偏移的鲁棒性与在线适应能力。
模仿学习中的分布偏移问题表现为:智能体无法为训练中未遇到的状态规划出合理动作。这主要源于专家示范在全环境上提供的状态-动作覆盖范围有限。本文提出一种鲁棒的离线到在线自适应模仿学习框架,以终身、多阶段方式应对分布偏移。离线阶段,利用判别器有效训练策略,结合补充演示数据扩展策略的状态-动作覆盖,提升对分布偏移的鲁棒性。在线推理阶段,框架可检测分布偏移,并基于在线经验进行自监督模仿学习,实现策略对在线环境的自适应。在MuJoCo环境中的大量实验表明,该方法比基线算法在分布偏移鲁棒性和在线环境适应性方面表现更优,验证了框架的有效性。
原文摘要 · Abstract (English)
Distribution shift in imitation learning refers to the problem that the agent cannot plan proper actions for a state that has not been visited during the training. This problem can be largely attributed to the inherently narrow state-action coverage provided by expert demonstrations over the full environment. In this paper, we propose a robust offline to adaptive online imitation learning framework that handles the distribution shift problem in a lifelong, multi-phase scheme. In the offline learning phase, we leverage supplementary demonstrations to broaden the state-action coverage of the policy by utilizing a discriminator to effectively train the policy with supplementary demonstrations, thereby enhancing the robustness of the policy to distribution shift. In the subsequent online inference phase, our framework detects the occurrence of distribution shift and conducts self-supervised imitation learning from online experiences to adapt the policy to the online environments. Through extensive evaluations in MuJoCo environments, we demonstrate that our method exhibits better robustness to distribution shift and better adaptation performance to online environments than the baseline algorithms, which indicates superior performance of our framework against the distribution shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。