用物体操作特性指导机器人抓取,提升泛化能力。
RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation

- 以操作特性为中间表示,分层生成任务计划并控制动作
- 在新任务上性能超现有方法50%以上,且对新环境鲁棒
- 可利用网络数据与低成本图像训练,无需额外机械臂数据
我们探索中间策略表示如何通过提供操作指导来促进泛化。现有表示如语言、目标图像和轨迹草图虽有帮助,但或信息不足,或过度指定导致策略不够鲁棒。本文提出基于操作特性的条件化策略,捕捉任务关键阶段的机器人位姿。操作特性具有表达力强且轻量的抽象特性,用户易标注,并可通过互联网大规模数据集高效迁移知识。所提方法RT-Affordance为分层模型:先根据任务语言生成操作特性计划,再以此条件化策略执行操作。该方法可灵活融合异构监督信号,包括大型网络数据集和机器人轨迹。此外,仅需收集低成本的领域内操作特性图像即可学习新任务,无需额外昂贵的机器人数据采集。实验表明,在多种新颖任务上,该方法性能超越现有方法超过50%,且在新场景下表现稳定。
原文摘要 · Abstract (English)
We explore how intermediate policy representations can facilitate generalization by providing guidance on how to perform manipulation tasks. Existing representations such as language, goal images, and trajectory sketches have been shown to be helpful, but these representations either do not provide enough context or provide over-specified context that yields less robust policies. We propose conditioning policies on affordances, which capture the pose of the robot at key stages of the task. Affordances offer expressive yet lightweight abstractions, are easy for users to specify, and facilitate efficient learning by transferring knowledge from large internet datasets. Our method, RT-Affordance, is a hierarchical model that first proposes an affordance plan given the task language, and then conditions the policy on this affordance plan to perform manipulation. Our model can flexibly bridge heterogeneous sources of supervision including large web datasets and robot trajectories. We additionally train our model on cheap-to-collect in-domain affordance images, allowing us to learn new tasks without collecting any additional costly robot trajectories. We show on a diverse set of novel tasks how RT-Affordance exceeds the performance of existing methods by over 50%, and we empirically demonstrate that affordances are robust to novel settings. Videos available at https://snasiriany.me/rt-affordance
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。