arXiv:2608.16222cs.ROcs.AI2026-08

构建600小时高精度人形动作与物体交互数据集,突破现有数据覆盖与质量瓶颈。

HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction

论文配图:HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction
图 1 · 摘自论文原文
  • 基于语言学框架FrameNet设计,系统覆盖人体动作与交互全貌
  • 提供亚毫米级追踪精度,支持全身运动与物体轨迹的高保真建模
  • 适用于人形机器人策略训练与计算机图形学动作先验建模

人形智能需学习复杂多样的全身动作与物理交互行为。现有具身数据集存在根本性局限:互联网视频缺乏精确物理状态与交互依据,实验室动作数据虽保真度高但行为覆盖狭窄。这一矛盾成为可扩展人形策略学习的关键瓶颈。本文提出HiPHI,一个规模超过600小时的高保真全身体动数据集,旨在系统最大化人类动作与交互流形的覆盖范围。该数据集基于语言学框架FrameNet设计,通过光学动捕系统实现全身亚毫米级空间标记追踪精度及物体网格级轨迹。我们进一步构建基准评估体系,涵盖动作空间多样性、交互真实性、物体一致性及物理人工智能应用。分析表明,相比现有数据集,HiPHI显著拓展了动作覆盖范围,同时保持高保真交互质量,为真实世界具身任务中人形策略的训练、评估与泛化提供了可扩展的数据基础,亦适用于计算机图形学中的动作先验模型构建。

原文摘要 · Abstract (English)

Humanoid intelligence requires learning over an extremely diverse space of whole-body motions and physically grounded interactions. However, existing embodied datasets remain fundamentally limited: internet-scale video data lack precise physical states and interaction grounding, while laboratory motion datasets provide high fidelity but only narrow behavioral coverage. This mismatch creates a critical bottleneck for scalable humanoid policy learning. We present HiPHI, a 600+ hour scale high-fidelity whole-body human motion dataset designed to systematically maximize coverage of the human motion and interaction manifold. HiPHI is theoretically guided by FrameNet, a linguistic framework organizing human primitives. Created using an optical motion capture pipeline, HiPHI provides sub-millimeter spatial marker tracking accuracy for full-body human motion and mesh-level object trajectories. We further introduce a benchmark suite evaluating motion-space diversity, interaction grounding, object consistency, and physical AI applications. Our analyses demonstrate that HiPHI significantly expands motion coverage compared to existing motion datasets while maintaining high-fidelity interaction quality, and establishes a scalable data foundation for training, evaluating, and generalizing humanoid policies in real-world embodied tasks, where similar extensions are also applicable to motion prior models in computer graphics.

动作捕捉人形机器人高精度数据物理交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。