arXiv:2507.15597cs.CVcs.LG2025-07被引 108

用真人视频训练机器人,让模型学会精细手部动作和指令理解。

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

  • 基于真人视频数据,采用物理指令微调提升动作精度。
  • 手部动作重建达毫米级准确,支持复杂操作任务学习。
  • 适合机器人抓取、操作等需要高灵巧性的实际应用。

我们提出 Being-H0,一种基于大规模真人视频训练的灵巧视觉-语言-动作模型(VLA)。现有 VLA 在需要高灵巧性的复杂操作任务中表现不佳,主要因依赖存在显著仿真到现实差距的合成数据或规模与多样性不足的遥操作演示。为突破数据瓶颈,我们利用人类双手作为基础操作器,挖掘网络数据中丰富的灵巧性与可扩展性。方法核心为物理指令微调,结合大规模真人视频预训练、三维空间对齐以支持3D推理,以及针对机器人任务的后训练适配。此外,提出部件级运动分词方法,实现毫米级动作轨迹重建,精准建模动作学习。为支持该范式,构建了整合运动捕捉、虚拟现实与仅RGB视频的综合数据清洗流程,形成包含数百万个基于动作的指令实例的大规模数据集。实验证明 Being-H0 在手部动作生成与指令遵循方面表现优异,且模型与数据规模均具良好扩展性。重要的是,应用物理指令微调后,其在真实机器人操作中展现出预期性能提升。

原文摘要 · Abstract (English)

We introduce Being-H0, a dexterous Vision-Language-Action model (VLA) trained on large-scale human videos. Existing VLAs struggle with complex manipulation tasks requiring high dexterity and generalize poorly to novel scenarios and tasks, primarily due to their reliance on synthetic data with significant sim-to-real gaps or teleoperated demonstrations lacking scale and diversity. To address this data bottleneck, we propose leveraging human hands as a foundation manipulator, capitalizing on the rich dexterity and scalability present in web data. Our approach centers on physical instruction tuning, a novel training paradigm that combines large-scale VLA pretraining from human videos, physical space alignment for 3D reasoning, and post-training adaptation for robotic tasks. Additionally, we introduce a part-level motion tokenization method which achieves millimeter-level reconstruction accuracy to model precise hand trajectories for action learning. To support our proposed paradigm, we further develop a comprehensive data curation pipeline that integrates heterogeneous sources -- including motion capture, VR, and RGB-only videos -- into a large-scale dataset with millions of motion-based instructional instances. We empirically show the excellence of Being-H0 in hand motion generation and instruction following, and it also scales well with model and data sizes. Importantly, we observe the expected gains of Being-H0 in real-world robotic manipulation as physical instruction tuning is applied. More details are available at https://beingbeyond.github.io/Being-H0.

视觉语言动作机器人操作真人数据灵巧操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。