从单目视频中高效重建人与物体交互的4D运动,数据量大且可扩展。
Efficient and Scalable Monocular Human-Object Interaction Motion Reconstruction
- 提出稀疏接触标注范式,结合多模态预测器实现高效标注
- 构建4DHOISolver优化框架,保证时空一致性和物理合理性
- 推出包含135类物体和133种动作的Open4DHOI数据集,适合机器人学习
通用机器人需从多样、大规模的人-物交互(HOI)中学习,以在真实世界中稳健运行。单目网络视频提供了几乎无限且易获取的数据源,涵盖前所未有的人类活动、物体与环境多样性。然而,从这些自然场景视频中准确、可扩展地提取4D交互数据仍是一大未解难题。为突破标注瓶颈,我们提出一种高效的稀疏接触标注范式,并开发InterPoint多模态预测器,驱动人机协同数据引擎。基于这些高效获取的标注,我们提出4DHOISolver优化框架,约束病态的4D HOI重建问题,保持高时空一致性与物理合理性。在此基础上,我们构建了开放的大型4D HOI数据集Open4DHOI,包含135种物体类型和133种动作。此外,我们验证了重建结果的有效性:利用强化学习代理成功模仿恢复的运动。数据与代码将公开于https://github.com/wenboran2002/open4dhoi_code。
原文摘要 · Abstract (English)
Generalized robots must learn from diverse, large-scale human-object interactions (HOI) to operate robustly in the real world. Monocular internet videos offer a nearly limitless and readily available source of data, capturing an unparalleled diversity of human activities, objects, and environments. However, accurately and scalably extracting 4D interaction data from these in-the-wild videos remains a significant and unsolved challenge. To overcome the annotation bottleneck, we introduce an efficient sparse contact annotation paradigm. To scale this process, we develop InterPoint, a multi-modal predictor that drives a human-in-the-loop data engine. Building upon these efficiently acquired annotations, we introduce 4DHOISolver, a novel optimization framework that constrains the ill-posed 4D HOI reconstruction problem, maintaining high spatio-temporal coherence and physical plausibility. Leveraging this framework, we introduce Open4DHOI, a new large-scale 4D HOI dataset featuring a diverse catalog of 135 object types and 133 actions. Furthermore, we demonstrate the effectiveness of our reconstructions by enabling an RL-based agent to imitate the recovered motions. Data and code will be publicly available at https://github.com/wenboran2002/open4dhoi_code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。