arXiv:2508.05506cs.CV2025-08ICCV被引 15

用3D先验提升单目视频中手物交互的重建精度

MagicHOI: Leveraging 3D Priors for Accurate Hand-object Reconstruction from Short Monocular Video Clips

  • 利用扩散模型生成新视角图像作为物体形状先验
  • 在视角受限时仍能准确重建被遮挡的物体部分
  • 适合真实场景下手物交互的三维重建任务

大多数基于RGB的手物重建方法依赖物体模板,而无模板方法通常假设物体完全可见。这一假设在实际场景中常不成立,固定相机视角和静态抓握导致物体部分区域不可见,造成不合理重建。为此,我们提出MagicHOI,一种从短单目交互视频中重建手与物体的方法,即使在视角变化有限的情况下也能工作。核心洞察是:尽管缺乏成对的3D手物数据,大规模新视角合成扩散模型可提供丰富的物体监督信息。该监督作为先验,在手物交互过程中正则化未观测到的物体区域。基于此,我们将新视角合成模型融入手物重建框架,并通过可见接触约束实现手物对齐。实验表明,MagicHOI显著优于现有最先进方法。同时证明,新视角合成扩散先验能有效正则化未观测区域,提升3D手物重建质量。

原文摘要 · Abstract (English)

Most RGB-based hand-object reconstruction methods rely on object templates, while template-free methods typically assume full object visibility. This assumption often breaks in real-world settings, where fixed camera viewpoints and static grips leave parts of the object unobserved, resulting in implausible reconstructions. To overcome this, we present MagicHOI, a method for reconstructing hands and objects from short monocular interaction videos, even under limited viewpoint variation. Our key insight is that, despite the scarcity of paired 3D hand-object data, large-scale novel view synthesis diffusion models offer rich object supervision. This supervision serves as a prior to regularize unseen object regions during hand interactions. Leveraging this insight, we integrate a novel view synthesis model into our hand-object reconstruction framework. We further align hand to object by incorporating visible contact constraints. Our results demonstrate that MagicHOI significantly outperforms existing state-of-the-art hand-object reconstruction methods. We also show that novel view synthesis diffusion priors effectively regularize unseen object regions, enhancing 3D hand-object reconstruction.

3D重建手物交互扩散模型单目视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。