用人类视频中的手部动作代替机器人数据,训练更适配机械臂控制的视觉模型。
Contrastive Action-Image Pre-training for Visuomotor Control

- 用人类第一视角视频中的3D手部关键点作为动作信号,构建对比学习框架。
- 仅用88小时机器人数据+3.2万小时人类视频,性能超DINOv2等主流模型。
- 在复杂抓取任务中提升30%以上,适合需要精细操作的机器人应用。
现有机器人视觉编码器面临根本瓶颈:缺乏大规模机器人数据用于预训练。以往工作通过互联网图像与语言数据或人类第一视角视频来弥补数据不足,但这些方法未学习视觉与动作的配对关系,而下游视觉-运动控制策略正依赖此类信号。尽管机器人轨迹是直接的配对数据源,却难以获取大规模样本,因此我们提出从海量人类视频中提取动作信号。本文引入CAIP(对比动作-图像预训练)方法,将大规模第一视角视频中的人类手部姿态作为末端执行器动作的代理。通过提取3D手部关键点——一种与机器人动作空间自然对齐的表示——并采用对比目标学习统一的动作-图像表征。利用32,041小时第一视角人类视频和仅88小时机器人操作数据,CAIP在包括Dexmate Vega和Sharpa Wave机械手在内的复杂现实世界灵巧操作任务上,显著优于DINOv2、SigLIP、MVP和R3M等先进视觉编码器,在折叠、倒液及精细操作任务中性能提升超过30%。结果表明,基于对比动作中心的预训练为获得更适合物理交互的鲁棒视觉表征提供了可扩展路径。
原文摘要 · Abstract (English)
Existing vision encoders for robotics face a fundamental bottleneck: robotic datasets lack the scale necessary for large-scale pre-training. Prior work circumvents this data scarcity by turning to internet-scale image and language data or egocentric human video. While these models show promise, neither paradigm learns from paired vision and action data, which downstream visuomotor control policies require. However, robot trajectories, the most direct source of this paired signal, are not available at pre-training scale, motivating us to extract action signals from abundant human video instead. To this end, we introduce CAIP (Contrastive Action-Image Pre-training), a vision encoder that treats human hand poses from large-scale egocentric video as a proxy for end-effector actions. By extracting 3D hand keypoints, a representation that aligns naturally with downstream robot action spaces, CAIP learns a unified action-image representation through a contrastive objective. Leveraging 32,041 hours of egocentric human video and only 88 hours of robotic manipulation data, CAIP outperforms state-of-the-art vision encoders including DINOv2, SigLIP, MVP, and R3M. Evaluated on a challenging real-world dexterous manipulation setup using Dexmate Vega and Sharpa Wave hands, CAIP yields performance gains of more than 30% on tasks involving folding, pouring, and fine-grained manipulation. Our results show that our method of contrastive action-centric pre-training yields a scalable path to achieving robust visual representations better suited for physical interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。