从自由动作视频中学习机器人抓取,仅需20次标注演示即可显著提升成功率。
3PoinTr: 3D Point Tracks for Learning Manipulation from Unconstrained Human Videos
- 通过预测3D点轨迹实现无需预设关键点的视频理解
- 真实场景下成功率比基线高25.0个百分点,仿真中高29.6个百分点
- 适合希望降低机器人示范成本的研究者与开发者
从人类视频中学习操作策略可大幅减少对昂贵机器人示范的需求,但现有方法通常依赖于受控动作、预定义关键点、人工标注或已知抓取位置等限制性假设。我们提出3PoinTr,一种从非约束人类视频中预训练样本高效机器人策略的方法,通过预测密集3D点轨迹实现。在非约束的人类示范视频中,人类可自由选择运动轨迹和操作策略,而非刻意模仿机器人动作。3PoinTr采用轻量级可见性感知变压器,学习场景点在视频中的运动规律,并训练一个闭环多任务机器人策略,灵活从预测的点轨迹中提取与动作相关先验。仅使用20次带动作标签的机器人示范,3PoinTr在真实世界任务中平均成功率比最强的行为克隆与视频预训练基线高出25.0个百分点,在仿真中高出29.6个百分点。针对性消融实验验证了关键设计选择的有效性,并证实从无动作标注视频中学习的价值。此外,3PoinTr的点轨迹预测变压器在部分遮挡点上保留监督信号方面优于强基线。
原文摘要 · Abstract (English)
Learning manipulation policies from human videos could greatly reduce the need for expensive robot demonstrations, but existing approaches typically require restrictive assumptions such as choreographed human motions, predefined keypoints, manual annotations, or known grasp locations. We propose 3PoinTr, a method for pretraining sample-efficient robot policies from unconstrained human videos by predicting dense 3D point tracks. In the unconstrained human demonstration videos, humans are free to follow whatever trajectories and manipulation strategies they see fit, rather than choreographing their motions to mimic a robot. 3PoinTr uses a lightweight visibility-aware transformer to learn how scene points should move from human videos, and then trains a closed-loop multitask robot policy to flexibly extract action-relevant priors from those predicted point tracks. With only 20 action-labeled robot demonstrations, 3PoinTr achieves a 25.0 percentage point higher average success rate than the strongest behavior cloning and video-pretraining baselines on real-world tasks, and a 29.6 percentage point higher average success rate in simulation. Targeted ablations support the key design choices and confirm the benefit of learning from actionless videos. We further show that 3PoinTr's point track prediction transformer outperforms a strong baseline by preserving supervision over partially occluded points. Project page: https://adamhung60.github.io/3PoinTr/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。