将真人视频转为机器人动作视频,实现大规模生成。
X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale
- 用生成模型将真人视频转为拟人机器人动作视频。
- 生成超360万帧机器人化视频帧,数据量达17小时合成+60小时真实视频。
- 用户评测显示动作连贯性与拟人度均优于现有方法。
具身智能的发展为智能拟人机器人带来了巨大潜力,但视觉-语言-动作(VLA)模型与世界模型的进步受限于大规模、多样化训练数据的缺乏。一种有前景的解决方案是“机器人化”网络规模的人类视频,已被证明对策略训练有效。然而,现有方法主要在第一视角视频上叠加机械臂,难以处理第三人称视频中的复杂全身动作与场景遮挡,不适用于人类视频的机器人化。为此,我们提出X-Humanoid,一种生成式视频编辑方法,将强大的Wan 2.2模型改造为视频到视频结构,并微调用于人体到拟人机器人的转换任务。该微调需要成对的人体-拟人机器人视频,因此我们设计了可扩展的数据生成流水线,利用Unreal Engine将社区资源转化为超过17小时的合成配对视频。随后,我们将训练好的模型应用于60小时的Ego-Exo4D视频,生成并发布了一个包含超过360万帧“机器人化”拟人视频的新大规模数据集。定量分析与用户研究证实本方法优于现有基线:69%的用户认为其运动一致性最佳,62.1%认为拟人正确性最优。
原文摘要 · Abstract (English)
The advancement of embodied AI has unlocked significant potential for intelligent humanoid robots. However, progress in both Vision-Language-Action (VLA) models and world models is severely hampered by the scarcity of large-scale, diverse training data. A promising solution is to "robotize" web-scale human videos, which has been proven effective for policy training. However, these solutions mainly "overlay" robot arms to egocentric videos, which cannot handle complex full-body motions and scene occlusions in third-person videos, making them unsuitable for robotizing humans. To bridge this gap, we introduce X-Humanoid, a generative video editing approach that adapts the powerful Wan 2.2 model into a video-to-video structure and finetunes it for the human-to-humanoid translation task. This finetuning requires paired human-humanoid videos, so we designed a scalable data creation pipeline, turning community assets into 17+ hours of paired synthetic videos using Unreal Engine. We then apply our trained model to 60 hours of the Ego-Exo4D videos, generating and releasing a new large-scale dataset of over 3.6 million "robotized" humanoid video frames. Quantitative analysis and user studies confirm our method's superiority over existing baselines: 69% of users rated it best for motion consistency, and 62.1% for embodiment correctness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。