构建百万小时人类行为视频数据集,推动具身智能发展
HumanNet: Scaling Human-centric Video Learning to One Million Hours

- 构建百万小时人本视角视频数据集,涵盖真实世界多种交互场景
- 1000小时人类第一人称视频训练效果超过100小时机器人实拍数据
- 适合研究具身智能、动作生成与人机迁移的学者和工程师
具身智能的发展越来越依赖可扩展的数据基础设施。尽管视觉和语言模型已借助互联网语料实现规模化,但物理交互学习仍受限于缺乏大规模、多样化且标注丰富的真人活动数据。本文提出 HumanNet,一个覆盖百万小时的人类行为视频语料库,涵盖第一人称与第三人称视角,包含精细动作、人-物交互、工具使用及长时序行为,覆盖多样真实环境。除原始视频外,还提供以交互为中心的标注信息,包括描述性字幕、运动说明以及手部与身体信号,支持运动感知与交互感知的学习。在规模之外,HumanNet引入系统化数据构建范式,将人本过滤、时间结构化、视角多样性与标注增强作为核心设计原则,使非结构化网络视频成为可扩展的表征学习、行为理解、运动生成与人到机器人的迁移基础。通过可控的视觉-语言-动作消融实验验证:在固定验证数据下,基于 Qwen VLM 模型继续训练 1000 小时来自 HumanNet 的第一人称视频,其表现优于使用 100 小时 Magic Cobot 机器人实拍数据的训练,表明第一人称人类视频可作为机器人数据的低成本高效替代方案。本工作旨在探索利用人类视频规模化构建具身基础模型的潜力,而非仅依赖机器人专属数据。
原文摘要 · Abstract (English)
Progress in embodied intelligence increasingly depends on scalable data infrastructure. While vision and language have scaled with internet corpora, learning physical interaction remains constrained by the lack of large, diverse, and richly annotated human activity data. We present HumanNet, a one-million-hour human-centric video corpus that captures how humans interact with the physical world at scale. HumanNet spans both first-person and third-person perspectives and covers fine-grained activities, human-object interactions, tool use, and long-horizon behaviors across diverse real-world environments. Beyond raw video, the dataset provides interaction-centric annotations, including captions, motion descriptions, and hand and body-related signals, enabling motion-aware and interaction-aware learning. Beyond scale, HumanNet introduces a systematic data curation paradigm for embodied learning, where human-centric filtering, temporal structuring, viewpoint diversity, and annotation enrichment are treated as first-class design principles. This design transforms unstructured internet video into a scalable substrate for representation learning, activity understanding, motion generation, and human-to-robot transfer. We conduct a first-step validation on the value of this design through controlled vision-language-action ablation: under a fixed set of validation data, continued training from the Qwen VLM model with 1000 hours of egocentric video drawn from HumanNet surpasses the continued training with 100 hours of real-robot data from Magic Cobot, indicating that egocentric human video could be a scalable and cost-effective substitute for robot data. By building this project, we aim to explore the opportunity to scale embodied foundation models using human-centric videos, rather than relying solely on robot-specific data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。