用视觉与触觉先验训练机器人,从人类演示视频中学习精准抓取动作。
OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer
- 基于物体中心思想,聚焦任务相关物体及其交互。
- 融合3D点云与百万级触觉图像数据,提升动作鲁棒性。
- 适合需要高精度操作的机器人教学场景。
我们提出OCRA,一种基于视频的人类到机器人动作迁移的物体中心框架,直接从人类示范视频中学习以实现稳健操作。该方法强调任务相关的物体及其相互作用,同时过滤无关背景,提供一种自然且可扩展的机器人教学方式。OCRA利用多视角RGB视频、先进的3D基础模型VGGT以及前沿的检测与分割模型,重建物体中心的3D点云,捕捉物体间的丰富交互。为应对仅靠视觉难以感知的属性,引入大规模触觉先验,基于超过一百万张触觉图像数据集。通过多模态模块ResFiLM融合3D与触觉先验,并输入扩散策略(Diffusion Policy)生成稳健的操作动作。在仅视觉和视听结合任务上的大量实验表明,OCRA显著优于现有基线与消融实验,验证了其从人类示范视频中学习的有效性。
原文摘要 · Abstract (English)
We present OCRA, an Object-Centric framework for video-based human-to-Robot Action transfer that learns directly from human demonstration videos to enable robust manipulation. Object-centric learning emphasizes task-relevant objects and their interactions while filtering out irrelevant background, providing a natural and scalable way to teach robots. OCRA leverages multi-view RGB videos, the state-of-the-art 3D foundation model VGGT, and advanced detection and segmentation models to reconstruct object-centric 3D point clouds, capturing rich interactions between objects. To handle properties not easily perceived by vision alone, we incorporate tactile priors via a large-scale dataset of over one million tactile images. These 3D and tactile priors are fused through a multimodal module (ResFiLM) and fed into a Diffusion Policy to generate robust manipulation actions. Extensive experiments on both vision-only and visuo-tactile tasks show that OCRA significantly outperforms existing baselines and ablations, demonstrating its effectiveness for learning from human demonstration videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。