arXiv:2412.05426cs.ROcs.AI2024-12被引 14

通过关键点实现机器人动作的高效通用模仿。

What's the Move? Hybrid Imitation Learning via Salient Points

  • 用点云和腕部图像识别任务关键点,分阶段控制机械臂动作。
  • 在440次真实实验中成功率86.7%,比最优基线高出41.1%。
  • 适合需要空间泛化与高精度操作的复杂抓取任务。

尽管模仿学习(IL)为教机器人执行多种行为提供了有前景的框架,但学习复杂任务仍具挑战性。现有IL策略在面对视觉与空间变化时泛化能力不足。本文提出SPHINX:基于关键点的混合模仿与执行,利用点云与腕部图像的多模态观测,结合低频稀疏路径点与高频密集末端执行器运动的混合动作空间。给定3D点云,SPHINX学习识别任务相关的点,即关键点,聚焦于语义有意义的特征以支持空间泛化。这些关键点作为锚点,用于预测远距离移动的路径点,如自由空间中的目标姿态。靠近关键点后,模型转而根据近距离腕部图像预测密集末端运动,完成任务的精细阶段。通过在不同操作阶段充分利用不同输入模态与动作表示的优势,SPHINX以样本高效、可泛化的方式应对复杂任务。该方法在4个真实世界与2个仿真任务中达到86.7%的成功率,平均优于当前最佳的IL基线41.1%(共440次真实试验)。此外,其在新视角、视觉干扰、空间布局及执行速度上均表现出良好泛化能力,相比最先进基线提速1.7倍。项目官网(http://sphinx-manip.github.io)提供数据采集、训练与评估的开源代码及补充视频。

原文摘要 · Abstract (English)

While imitation learning (IL) offers a promising framework for teaching robots various behaviors, learning complex tasks remains challenging. Existing IL policies struggle to generalize effectively across visual and spatial variations even for simple tasks. In this work, we introduce SPHINX: Salient Point-based Hybrid ImitatioN and eXecution, a flexible IL policy that leverages multimodal observations (point clouds and wrist images), along with a hybrid action space of low-frequency, sparse waypoints and high-frequency, dense end effector movements. Given 3D point cloud observations, SPHINX learns to infer task-relevant points within a point cloud, or salient points, which support spatial generalization by focusing on semantically meaningful features. These salient points serve as anchor points to predict waypoints for long-range movement, such as reaching target poses in free-space. Once near a salient point, SPHINX learns to switch to predicting dense end-effector movements given close-up wrist images for precise phases of a task. By exploiting the strengths of different input modalities and action representations for different manipulation phases, SPHINX tackles complex tasks in a sample-efficient, generalizable manner. Our method achieves 86.7% success across 4 real-world and 2 simulated tasks, outperforming the next best state-of-the-art IL baseline by 41.1% on average across 440 real world trials. SPHINX additionally generalizes to novel viewpoints, visual distractors, spatial arrangements, and execution speeds with a 1.7x speedup over the most competitive baseline. Our website (http://sphinx-manip.github.io) provides open-sourced code for data collection, training, and evaluation, along with supplementary videos.

模仿学习关键点机器人操控多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。