arXiv:2606.32009cs.RO2026-06

用真人视频生成可执行机器人动作,让高自由度人形机器人零样本学习。

Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments

论文配图:Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments
图 1 · 摘自论文原文
  • 通过同步第一/第三人称视频,将人类动作转为机器人可执行的60自由度动作
  • 在真实机器人上验证,仅用转换后的人类数据即可实现任务泛化
  • 相比传统遥操作,演示生成效率提升4.8至7.2倍,适合无机器人数据场景

视觉-语言-动作(VLA)模型在不同机器人形态上训练需高质量观测-动作监督,但高自由度人形机器人数据难以获取。遥操作提供控制器对齐监督,而人类第一人称视频虽涵盖多样双臂操作,却无法直接生成可执行动作。本文提出Human-as-Humanoid框架,通过联合对齐机器人本体、感知设置与动作标签接口,实现近实时的人类中心动作生成。基于具备60自由度上肢的人形模型PrimeU,该方法利用同步的第一/第三人称视频,将部署对齐的视角观测与第三人称运动恢复结合,通过分阶段逆运动学(IK)将恢复的人类动作重定向为控制器对齐的60自由度动作块,并使用前向运动学(FK)感知监督训练VLA模型,以保持手腕和指尖的任务空间几何结构。该流程将大规模人类演示从视觉观察转化为目标人形机器人的可执行观测-动作监督。实验在运动恢复、机器人动作空间及真实机器人部署层面验证了转换链的有效性。数据采集分析显示,相较于人形机器人遥操作,该方法提升4.8至7.2倍原始演示吞吐量;下游任务中,仅用转换后人类标签训练的策略在无需目标任务机器人演示的情况下即能泛化至真实机器人部署。项目主页:https://zgc-embodyai.github.io/Human-as-Humanoid。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models across robot embodiments require high-quality observation--action supervision to learn deployable action distributions, yet scaling such robot data remains difficult, especially for high-DoF humanoids. Teleoperation provides controller-aligned supervision, while human egocentric videos capture diverse bimanual manipulation but do not directly provide executable robot actions. We introduce Human-as-Humanoid, a human-to-humanoid supervision framework that enables near-real-time human-centric action generation, making human demonstrations usable for high-DoF humanoid VLA training by jointly aligning the robot embodiment, the sensing setup, and the action-label interface. Built on PrimeU, a human-aligned 60-DoF upper-body humanoid, Human-as-Humanoid uses synchronized ego-exo videos to pair deployment-aligned egocentric observations with exocentric motion recovery, retargets the recovered human motion through staged Inverse Kinematics (IK) into controller-aligned 60-DoF action chunks, and trains the VLA model with Forward Kinematics (FK)-aware supervision to preserve wrist and fingertip task-space geometry. This converts large-scale human demonstrations from visual observations into executable observation--action supervision for the target humanoid. Experiments validate the conversion chain at the motion-recovery, robot-action-space, and real-robot deployment levels. Human-as-Humanoid yields a 4.8--7.2x raw demonstration-throughput gain over humanoid teleoperation in our data-collection analysis, and on several downstream tasks, policies post-trained only with the converted human labels generalize to real-robot deployment without target-task robot demonstrations. The official project website is available at https://zgc-embodyai.github.io/Human-as-Humanoid.

人形机器人动作迁移视觉-语言-动作零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。