arXiv:2503.15898cs.CV2025-03CVPR被引 13

从单图重建真实场景中人与物的3D交互,构建首个开放词汇数据集。

Reconstructing In-the-Wild Open-Vocabulary Human-Object Interactions

  • 基于单图3D重建技术,构建可标注人物与物体3D交互的流水线。
  • 创建2500+个3D HOI资产,推出首个野外开放词汇3D HOI数据集Open3DHOI。
  • 提出Gaussian-HOI优化器,高效恢复空间交互与接触区域,适合3D理解研究者。

从单张图像重建人-物交互(HOI)是计算机视觉的基础任务。现有方法多在室内场景上训练与测试,受限于缺乏3D数据,尤其受物体多样性制约,难以泛化至真实复杂场景。以往3D HOI数据集的局限性主要源于3D物体资产获取困难。随着单图3D重建技术的发展,现可从2D HOI图像中重建多样物体。本文提出一套从单图标注精细3D人体、物体及其交互的流程,从已有2D HOI数据集中标注超过2500个3D HOI资产,并构建首个开放词汇的野外3D HOI数据集Open3DHOI,作为未来评测基准。此外,设计新型Gaussian-HOI优化器,实现人体与物体间空间交互的高效重建并学习接触区域。除3D HOI重建外,还提出多个新任务以推动后续研究。数据与代码将公开于https://wenboran2002.github.io/3dhoi。

原文摘要 · Abstract (English)

Reconstructing human-object interactions (HOI) from single images is fundamental in computer vision. Existing methods are primarily trained and tested on indoor scenes due to the lack of 3D data, particularly constrained by the object variety, making it challenging to generalize to real-world scenes with a wide range of objects. The limitations of previous 3D HOI datasets were primarily due to the difficulty in acquiring 3D object assets. However, with the development of 3D reconstruction from single images, recently it has become possible to reconstruct various objects from 2D HOI images. We therefore propose a pipeline for annotating fine-grained 3D humans, objects, and their interactions from single images. We annotated 2.5k+ 3D HOI assets from existing 2D HOI datasets and built the first open-vocabulary in-the-wild 3D HOI dataset Open3DHOI, to serve as a future test set. Moreover, we design a novel Gaussian-HOI optimizer, which efficiently reconstructs the spatial interactions between humans and objects while learning the contact regions. Besides the 3D HOI reconstruction, we also propose several new tasks for 3D HOI understanding to pave the way for future work. Data and code will be publicly available at https://wenboran2002.github.io/3dhoi.

3D HOI单图重建开放词汇交互理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。