让机器人快速学会抓捏会变形的物体,成功率超80%。
Rapid Adaptation of Particle Dynamics for Generalized Deformable Object Mobile Manipulation
- 用粒子位置变化捕捉形变,扩展快速运动适应方法
- 仿真中学习动态嵌入,真实世界用视觉和动作推断
- 适用于多种物体、不同变形特性,适合移动机械臂
针对未知动力学的可变形物体操作挑战,本文提出RAPiD方法。在非刚性物体中,动力学参数决定了其受力后的形变与运动方式,是成功完成操作任务的关键。已有方法如快速运动适应(RMA)可在腿式行走和刚体操作中处理未知动力学,通过仿真中监督学习将物体质量、位置等动力学编码为隐向量,并让策略根据该向量调整动作。然而,可变形物体的动力学不仅包含质量与位置,还包括形状变化。本文关键洞察是:仿真中的物体粒子真实位置能有效反映形状变化,从而可将RMA扩展至可变形物体。基于此,RAPiD采用两阶段方法:第一阶段在仿真中学习一个条件于物体动力学嵌入的视觉-运动策略(使用质量、粒子位置等特权信息);第二阶段学习仅用机器人视觉观测和动作来推断该嵌入,实现真实世界迁移。在22自由度移动机械臂上,该方法在两个基于视觉的可变形物体移动操作任务中均实现超过80%的成功率,且对不同物体类型、实例和动力学具有鲁棒性。
原文摘要 · Abstract (English)
We address the challenge of learning to manipulate deformable objects with unknown dynamics. In non-rigid objects, the dynamics parameters define how they react to interactions -- how they stretch, bend, compress, and move -- and they are critical to determining the optimal actions to perform a manipulation task successfully. In other robotic domains, such as legged locomotion and in-hand rigid object manipulation, state-of-the-art approaches can handle unknown dynamics using Rapid Motor Adaptation (RMA). Through a supervised procedure in simulation that encodes each rigid object's dynamics, such as mass and position, these approaches learn a policy that conditions actions on a vector of latent dynamic parameters inferred from sequences of state-actions. However, in deformable object manipulation, the object's dynamics not only includes its mass and position, but also how the shape of the object changes. Our key insight is that the recent ground-truth particle positions of a deformable object in simulation capture changes in the object's shape, making it possible to extend RMA to deformable object manipulation. This key insight allows us to develop RAPiD, a two-phase method that learns to perform real-robot deformable object mobile manipulation by: 1) learning a visuomotor policy conditioned on the object's dynamics embedding, which is encoded from the object's privileged information in simulation, such as its mass and ground-truth particle positions, and 2) learning to infer this embedding using non-privileged information instead, such as robot visual observations and actions, so that the learned policy can transfer to the real world. On a mobile manipulator with 22 degrees of freedom, RAPiD enables over 80%+ success rates across two vision-based deformable object mobile manipulation tasks in the real world, under various object dynamics, categories, and instances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。