用运动预测和阻抗控制实现通用铰接物体抓取的零样本迁移。
Watch Less, Feel More: Sim-to-Real RL for Generalizable Articulated Object Manipulation via Motion Adaptation and Impedance Control
- 基于历史观测预测物体运动与属性,结合可变阻抗控制减少对视觉的依赖。
- 在真实世界中对多种未见物体实现84%成功率,首次达成零样本转移。
- 适合需要高泛化性、低视觉依赖的机器人操控场景。
与刚体操作相比,铰接物体操作的独特挑战在于物体自身构成动态环境。本文提出一种基于强化学习的新型流水线,结合可变阻抗控制与运动适应机制,利用观测历史实现可泛化的铰接物体操作,重点提升零样本仿真到现实的平滑性和灵巧性。为缩小仿真与现实之间的差距,该流水线不直接将视觉特征(如RGBD/点云)作为策略输入,而是先通过现成模块提取低维有用数据。同时,通过观测历史推断物体运动及其内在属性,并在仿真与真实世界中均采用阻抗控制。此外,我们设计了具备强随机化和专用奖励机制(任务感知与运动感知)的训练设置,实现多阶段端到端操作,无需启发式运动规划。据我们所知,该策略是首个在大量不同未见物体上实现在真实世界84%成功率的方案。
原文摘要 · Abstract (English)
Articulated object manipulation poses a unique challenge compared to rigid object manipulation as the object itself represents a dynamic environment. In this work, we present a novel RL-based pipeline equipped with variable impedance control and motion adaptation leveraging observation history for generalizable articulated object manipulation, focusing on smooth and dexterous motion during zero-shot sim-to-real transfer. To mitigate the sim-to-real gap, our pipeline diminishes reliance on vision by not leveraging the vision data feature (RGBD/pointcloud) directly as policy input but rather extracting useful low-dimensional data first via off-the-shelf modules. Additionally, we experience less sim-to-real gap by inferring object motion and its intrinsic properties via observation history as well as utilizing impedance control both in the simulation and in the real world. Furthermore, we develop a well-designed training setting with great randomization and a specialized reward system (task-aware and motion-aware) that enables multi-staged, end-to-end manipulation without heuristic motion planning. To the best of our knowledge, our policy is the first to report 84\% success rate in the real world via extensive experiments with various unseen objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。