从人类操作视频中自动提取密集动作轨迹,提升机器人学习数据效率
Learning from Watching: Scalable Extraction of Manipulation Trajectories from Human Videos
- 结合视觉大模型与点跟踪技术,追踪操作过程中的关键点轨迹
- 在全过程中准确追踪关键点,支持大规模数据采集
- 适合需要海量演示数据的机器人技能学习研究者
训练大规模机器人模型通常依赖真实机器人平台收集高质量数据,成本高昂且耗时,无论是通过远程操控还是脚本演示。为实现数据规模化,许多研究转向利用网络上可获取的人类操作视频。然而,现有方法多集中于手部检测或物体位姿估计,未能充分挖掘视频中丰富的交互线索。本文提出一种新方法,结合大型基础模型的视频理解能力与点跟踪技术,从人类操作视频中提取任务相关关键点的密集轨迹,实现对互联网规模演示视频的更全面利用。实验表明,该方法能精准追踪整个操作过程中的关键点,为更高效、可扩展的机器人学习铺平道路。
原文摘要 · Abstract (English)
Collecting high-quality data for training large-scale robotic models typically relies on real robot platforms, which is labor-intensive and costly, whether via teleoperation or scripted demonstrations. To scale data collection, many researchers have turned to leveraging human manipulation videos available online. However, current methods predominantly focus on hand detection or object pose estimation, failing to fully exploit the rich interaction cues embedded in these videos. In this work, we propose a novel approach that combines large foundation models for video understanding with point tracking techniques to extract dense trajectories of all task-relevant keypoints during manipulation. This enables more comprehensive utilization of Internet-scale human demonstration videos. Experimental results demonstrate that our method can accurately track keypoints throughout the entire manipulation process, paving the way for more scalable and data-efficient robot learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。