用4D视频表征预训练机器人模型,让人类动作数据直接赋能机器人控制。
Pre-training Auto-regressive Robotic Models with 4D Representations

- 通过单目深度估计将2D视频转为3D点追踪,构建时序4D表征。
- 在多种机器人环境中,性能显著优于基线方法。
- 适合做通用机器人控制的预训练,尤其擅长从人类视频迁移。
基础模型在大规模无标签数据上预训练后,在自然语言和计算机视觉领域展现出卓越的泛化能力,凸显了预训练的重要性。然而,机器人领域尚未取得类似成功,受限于高昂的机器人标注成本或缺乏能有效建模物理世界的表征。本文提出ARM4R——一种自回归机器人模型,利用从人类视频数据中学习到的低层4D表征,实现更优的机器人预训练。具体而言,我们聚焦于通过时间序列单目深度估计将2D表征升维至3D空间,得到3D点追踪表征,形成4D表示。这些4D表征与机器人状态表示之间保持共享几何结构(至线性变换),从而实现从人类视频数据到低层机器人控制的高效迁移。实验表明,ARM4R能高效从人类视频数据迁移至机器人任务,并在多种机器人环境与配置下持续提升性能。
原文摘要 · Abstract (English)
Foundation models pre-trained on massive unlabeled datasets have revolutionized natural language and computer vision, exhibiting remarkable generalization capabilities, thus highlighting the importance of pre-training. Yet, efforts in robotics have struggled to achieve similar success, limited by either the need for costly robotic annotations or the lack of representations that effectively model the physical world. In this paper, we introduce ARM4R, an Auto-regressive Robotic Model that leverages low-level 4D Representations learned from human video data to yield a better pre-trained robotic model. Specifically, we focus on utilizing 3D point tracking representations from videos derived by lifting 2D representations into 3D space via monocular depth estimation across time. These 4D representations maintain a shared geometric structure between the points and robot state representations up to a linear transformation, enabling efficient transfer learning from human video data to low-level robotic control. Our experiments show that ARM4R can transfer efficiently from human video data to robotics and consistently improves performance on tasks across various robot environments and configurations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。