arXiv:2411.03294cs.ROcs.AI2024-11被引 8

让机器人在陌生环境中自动恢复,提升模仿学习的鲁棒性。

Out-of-Distribution Recovery with Object-Centric Keypoint Inverse Policy for Visuomotor Imitation Learning

  • 基于物体关键点梯度构建逆向恢复策略
  • 真实机器人实验中OOD表现提升77.7%
  • 无需额外数据,可自动生成新演示用于持续学习

我们提出一种以物体为中心的分布外恢复(OCR)框架,解决视觉运动策略学习中分布外(OOD)场景的挑战。传统行为克隆(BC)方法严重依赖大量标注数据覆盖,难以应对陌生空间状态。我们的方法不依赖额外数据采集,而是从原始训练数据中的物体关键点流形梯度推断出逆向策略,构建恢复策略。该策略可作为通用插件添加至任意基础视觉运动BC策略,无需特定方法适配,能引导系统返回训练分布,确保在分布外情况下的任务成功率。我们在仿真和真实机器人实验中验证了该框架的有效性,在分布外场景下相较基线策略提升77.7%。此外,我们展示了OCR自主收集演示以支持持续学习的能力。整体上,该框架为提升真实场景下视觉运动策略的鲁棒性迈出了重要一步。

原文摘要 · Abstract (English)

We propose an object-centric recovery (OCR) framework to address the challenges of out-of-distribution (OOD) scenarios in visuomotor policy learning. Previous behavior cloning (BC) methods rely heavily on a large amount of labeled data coverage, failing in unfamiliar spatial states. Without relying on extra data collection, our approach learns a recovery policy constructed by an inverse policy inferred from the object keypoint manifold gradient in the original training data. The recovery policy serves as a simple add-on to any base visuomotor BC policy, agnostic to a specific method, guiding the system back towards the training distribution to ensure task success even in OOD situations. We demonstrate the effectiveness of our object-centric framework in both simulation and real robot experiments, achieving an improvement of 77.7\% over the base policy in OOD. Furthermore, we show OCR's capacity to autonomously collect demonstrations for continual learning. Overall, we believe this framework represents a step toward improving the robustness of visuomotor policies in real-world settings.

机器人模仿学习鲁棒性恢复策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。