arXiv:2606.22136cs.RO2026-06被引 2

用生成模型合成真人手部操作数据,提升机器人抓取泛化能力。

Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data

论文配图:Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data
图 1 · 摘自论文原文
  • 基于语言、物体和场景生成真人视角操作视频
  • 在18项任务中零样本成功率从8.3%提升至38.9%
  • 适合想提升机器人灵巧操作泛化能力的研究者

灵巧操作的泛化需覆盖多种物体、场景和任务,但现有数据存在规模与对齐性的权衡:远程操控数据与机器人部署对齐但收集成本高;仿真可扩展但存在仿真到现实的差距;真实第一人称视频虽规模大,却与机器人部署不匹配。本文提出Wh0框架,利用生成式视频世界模型作为可扩展且可控的第一人称人类手部操作数据源,以激活预训练灵巧操作视觉-语言-动作(VLA)模型的能力。在语言、物体和场景条件下,Wh0通过生成式世界模型构建了包含5万条轨迹的WM-H数据集。随后,通过手部运动重建与视觉编辑,将生成视频转化为机器人可训练的监督信号。结合少量真实机器人数据进行联合训练后,预训练的VLA模型能有效适配灵巧操作部署。在18个真实世界灵巧操作任务中,相比仅在机器人数据上微调的模型,Wh0使未见过任务的零样本成功率从8.3%提升至38.9%。消融实验进一步表明,可扩展生成与场景/身体对齐是性能提升的关键因素。视频与开源代码见项目网站:https://chenyt31.github.io/wh0.github.io/

原文摘要 · Abstract (English)

Scaling dexterous manipulation requires generalization across objects, scenes, and tasks, yet existing data sources face a trade-off between scale and scene/embodiment alignment: teleoperation data is well aligned with robot deployment but expensive to collect; simulation is scalable but limited by the sim-to-real gap; and real egocentric videos scale effectively but remain misaligned with robot deployment. We propose Wh0, a framework that uses generative video world models as scalable and controllable sources of egocentric human-hand manipulation data to unlock the manipulation capabilities of pretrained dexterous VLA models. Conditioned on language, objects, and scenes, Wh0 uses a generative world model to produce WM-H, a 50k-episode dataset of egocentric human-object interaction videos. Wh0 then converts the generated videos into robot-trainable supervision through hand motion reconstruction and visual editing. Co-trained with a limited amount of real robot data, WM-H adapts pretrained VLA models to dexterous manipulation deployment. Across 18 real-world dexterous manipulation tasks, compared with a model post-trained only on robot data, Wh0 improves zero-shot success on unseen tasks from 8.3% to 38.9%. Ablation studies further show that scalable generation and scene/embodiment alignment are key drivers of performance gains. Videos and open-source code can be found on our project website: https://chenyt31.github.io/wh0.github.io/.

生成模型灵巧操作数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。