arXiv:2410.19379cs.RO2024-10中稿 · IEEE CASE 2024被引 2

用动态状态监督提升视觉模仿学习的非抓取操作性能

Visual Imitation Learning of Non-Prehensile Manipulation Tasks with Dynamics-Supervised Models

  • 通过直接预测物体位置速度加速度,构建更懂物理动力学的世界模型
  • 预训练后成功率达85%,比原方法提升64个百分点
  • 适合需要强泛化能力的复杂动态任务研究者

与拾取放置等准静态任务不同,非抓取操作等动态任务对基于视觉的控制更具挑战性,关键在于提取任务相关特征。传统视觉模仿学习通过反向传播策略损失来学习特征,但泛化能力有限。相比之下,世界模型可学习更具通用性的视觉骨干。现有方法仅预测下一帧RGB图像,难以充分捕捉任务相关的动力学信息。本文提出直接监督目标动态状态(动态映射)的方法,让世界模型同时预测环境刚体的位置、速度和加速度。在两个非预设二维任务(平衡-到达、倒桶)中验证,无论解耦、联合或端到端训练,还是前馈或循环策略架构,动态映射均显著提升性能。尤其在预训练阶段,成功率从21%跃升至85%。虽冻结后的动态模型在同域任务中表现良好,但在跨域任务中泛化能力下降。

原文摘要 · Abstract (English)

Unlike quasi-static robotic manipulation tasks like pick-and-place, dynamic tasks such as non-prehensile manipulation pose greater challenges, especially for vision-based control. Successful control requires the extraction of features relevant to the target task. In visual imitation learning settings, these features can be learnt by backpropagating the policy loss through the vision backbone. Yet, this approach tends to learn task-specific features with limited generalizability. Alternatively, learning world models can realize more generalizable vision backbones. Utilizing the learnt features, task-specific policies are subsequently trained. Commonly, these models are trained solely to predict the next RGB state from the current state and action taken. But only-RGB prediction might not fully-capture the task-relevant dynamics. In this work, we hypothesize that direct supervision of target dynamic states (Dynamics Mapping) can learn better dynamics-informed world models. Beside the next RGB reconstruction, the world model is also trained to directly predict position, velocity, and acceleration of environment rigid bodies. To verify our hypothesis, we designed a non-prehensile 2D environment tailored to two tasks: "Balance-Reaching" and "Bin-Dropping". When trained on the first task, dynamics mapping enhanced the task performance under different training configurations (Decoupled, Joint, End-to-End) and policy architectures (Feedforward, Recurrent). Notably, its most significant impact was for world model pretraining boosting the success rate from 21% to 85%. Although frozen dynamics-informed world models could generalize well to a task with in-domain dynamics, but poorly to a one with out-of-domain dynamics.

视觉模仿动态建模非抓取操作世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。