arXiv:2606.17200cs.RO2026-06被引 2

用人类视角视频生成机器人可学的伪动作数据,提升智能体预训练效果

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining

论文配图:ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining
图 1 · 摘自论文原文
  • 将人类第一视角视频转为机器人格式的伪动作轨迹,统一不同数据源
  • 在4530小时机器人数据+1480小时人类伪动作数据上训练,性能领先
  • 适合做具身智能、多模态预训练和真实世界操作任务的研究者

视觉-语言-动作(VLA)模型受益于大规模多样化的具身数据,但机器人轨迹采集成本高且耗时。近期研究表明,大规模第一人称人类视频可为预训练提供互补的真实世界监督。然而,由于动作空间、身体结构、时间动态和监督质量差异,联合训练人类与机器人数据仍具挑战。本文提出ACE-Ego-0,一个统一的VLA预训练框架,融合异构数据源。为从第一人称人类视频中提取大规模预训练监督,构建可扩展的视频到动作流水线,将原始人类视频转化为机器人格式的伪动作轨迹。为使这些标签与机器人示范可比,采用基于相机空间动作、形态条件和时间对齐动作分块的统一动作表示。为鲁棒利用含噪声的伪动作监督,设计可靠性感知训练目标,并引入人类辅助损失,聚焦可靠信号。在4.53K小时机器人与仿真数据,以及1.48K小时伪动作标注的人类视频数据上实例化ACE-Ego-0。实验表明,在可靠性加权下引入大规模人类监督,持续提升联合预训练与监督微调性能。在RoboCasa GR1 TableTop与RoboTwin 2.0上达到当前最优表现,并展示出强大的真实双臂操作迁移能力。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models benefit from large-scale and diverse embodied data, yet scaling robot trajectory collection is costly and labor-intensive. Recent advances show that large-scale egocentric human videos provide complementary real-world supervision in pretraining. However, joint training on human and robot data remains challenging due to divergences in action spaces, embodiment structures, temporal dynamics, and supervision quality. We introduce ACE-EGO-0, a unified VLA pretraining framework jointly leveraging heterogeneous data sources. To extract large-scale pretraining supervision from egocentric human videos, we build a scalable egocentric video-to-action pipeline that converts raw human videos into robot-format pseudo-action trajectories. To make these labels comparable with robot demonstrations, ACE-EGO-0 uses a unified action representation based on camera-space actions, morphology conditioning, and time-aligned action chunking. To robustly leverage noisy pseudo-action supervision from egocentric human videos, we formulate a reliability-aware training objective with a human auxiliary loss that concentrates supervision on reliable signals. We instantiate ACE-EGO-0 on 4.53K hours of robot and simulation data, together with 1.48K hours of pseudo-action-labeled egocentric human data. Experiments show that incorporating large-scale human supervision under reliability-aware weighting consistently improves both unified joint pretraining and supervised fine-tuning. ACE-EGO-0 achieves state-of-the-art performance on RoboCasa GR1 TableTop and RoboTwin 2.0, while demonstrating strong transfer to real-world bimanual manipulation.

具身智能视觉语言动作数据融合动作建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。