arXiv:2602.16710cs.RO2026-02被引 77

用2万小时人类第一视角视频训练机器人,让机械手动作更灵活可靠。

EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data

  • 基于20854小时人类第一视角视频训练视觉语言动作模型
  • 数据量越大,模型误差越小,且与真实机器人表现强相关
  • 只需少量机器人数据就能快速适配新任务,适合高自由度机械手

人类行为是学习物理智能最具可扩展性的数据来源之一,但如何有效利用其进行灵巧操作仍不明确。现有工作仅在受限场景中实现从人到机器人的迁移,尚不清楚大规模人类数据能否支持高自由度、精细的灵巧操作。本文提出EgoScale,一个基于大规模第一视角人类数据的灵巧操作迁移框架。在超过20,854小时带动作标注的第一视角人类视频上训练视觉语言动作(VLA)模型,数据量是以往工作的20倍以上,并发现人类数据规模与验证损失之间存在对数线性缩放规律。该验证损失与下游真实机器人性能高度相关,确立了大规模人类数据作为可预测监督信号的地位。除规模外,提出简单两阶段迁移方案:先进行大规模人类预训练,再进行轻量级人类-机器人对齐微调。该方法实现了强长时程灵巧操作和一次示例任务适应,仅需极少机器人监督。最终策略在22自由度灵巧机械手上相比无预训练基线提升平均成功率54%,并能有效迁移到低自由度机械手,表明大规模人类运动提供了一种可复用、与具体身体形态无关的动作先验。

原文摘要 · Abstract (English)

Human behavior is among the most scalable sources of data for learning physical intelligence, yet how to effectively leverage it for dexterous manipulation remains unclear. While prior work demonstrates human to robot transfer in constrained settings, it is unclear whether large scale human data can support fine grained, high degree of freedom dexterous manipulation. We present EgoScale, a human to dexterous manipulation transfer framework built on large scale egocentric human data. We train a Vision Language Action (VLA) model on over 20,854 hours of action labeled egocentric human video, more than 20 times larger than prior efforts, and uncover a log linear scaling law between human data scale and validation loss. This validation loss strongly correlates with downstream real robot performance, establishing large scale human data as a predictable supervision source. Beyond scale, we introduce a simple two stage transfer recipe: large scale human pretraining followed by lightweight aligned human robot mid training. This enables strong long horizon dexterous manipulation and one shot task adaptation with minimal robot supervision. Our final policy improves average success rate by 54% over a no pretraining baseline using a 22 DoF dexterous robotic hand, and transfers effectively to robots with lower DoF hands, indicating that large scale human motion provides a reusable, embodiment agnostic motor prior.

灵巧操作第一视角大模型迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。