arXiv:2504.04612cs.ROcs.AI2025-04中稿 · CoRL被引 13

通过观察人类用工具视频,让机器人学会通用工具操作技能。

Tool-as-Interface: Learning Robot Policies from Observing Human Tool Use

  • 用双摄像头重建3D场景并合成新视角,提升视角鲁棒性。
  • 采用工具中心化动作和分割观测,减少人机差异影响。
  • 比遥操作扩散策略成功率高71%,数据采集时间减少超四成。

工具使用是机器人执行复杂现实任务的关键,但学习此类技能需要大量数据。尽管遥操作广泛使用,但其速度慢、延迟敏感,且不适用于动态任务。相比之下,人类视频无需特殊硬件即可自然收集数据,但因视角变化和人机差异给机器人学习带来挑战。为此,我们提出一种从人类工具使用中迁移知识的框架。为增强策略对视角变化的鲁棒性,使用两个RGB相机重建3D场景,并利用高斯点云进行新视角合成。通过分割观测与工具中心化、任务空间的动作设计,实现与身体无关的视觉-运动策略学习。在多样化的工具使用任务中,所学策略展现出强泛化能力,对人类扰动、相机运动和机器人基座移动均具有鲁棒性。相比基于遥操作的扩散策略,任务成功率提升71%;数据采集时间分别减少77%(相对于遥操作)和41%(相对于当前最优接口)。

原文摘要 · Abstract (English)

Tool use is essential for enabling robots to perform complex real-world tasks, but learning such skills requires extensive datasets. While teleoperation is widely used, it is slow, delay-sensitive, and poorly suited for dynamic tasks. In contrast, human videos provide a natural way for data collection without specialized hardware, though they pose challenges on robot learning due to viewpoint variations and embodiment gaps. To address these challenges, we propose a framework that transfers tool-use knowledge from humans to robots. To improve the policy's robustness to viewpoint variations, we use two RGB cameras to reconstruct 3D scenes and apply Gaussian splatting for novel view synthesis. We reduce the embodiment gap using segmented observations and tool-centric, task-space actions to achieve embodiment-invariant visuomotor policy learning. We demonstrate our framework's effectiveness across a diverse suite of tool-use tasks, where our learned policy shows strong generalization and robustness to human perturbations, camera motion, and robot base movement. Our method achieves a 71\% improvement in task success over teleoperation-based diffusion policies and dramatically reduces data collection time by 77\% and 41\% compared to teleoperation and the state-of-the-art interface, respectively.

机器人工具使用视觉学习泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。