用连续接触信号提升单图3D人物交互重建精度
LEXIS: LatEnt ProXimal Interaction Signatures for 3D HOI from an Image

- 提出InterFields表示全身与物体的密集连续接近关系
- 在Open3DHOI和BEHAVE上接触与形变误差降低18%以上
- 适合需要高真实感3D场景理解的研究者使用
从单张RGB图像重建3D人-物交互对感知系统至关重要,但需捕捉身体与物体间的细微物理耦合。现有方法依赖稀疏二值接触线索,无法建模自然交互中的连续邻近与密集空间关系。本文提出InterFields,一种编码全身与物体表面密集连续接近关系的表示。由于单图推断该场固有病态,我们基于动作与物体几何结构的典型模式,设计了LEXIS——通过VQ-VAE学习的离散交互签名流形。进一步构建LEXIS-Flow扩散框架,利用这些签名联合估计人体与物体网格及其InterFields。InterFields实现引导式精炼,无需后处理优化即可生成物理合理、邻近感知的重建结果。在Open3DHOI与BEHAVE数据集上,本方法显著优于当前最先进基线,在重建、接触与邻近质量方面均有提升。方法兼具更强泛化能力,生成结果被评估为更具真实感,推动迈向整体3D场景理解。代码与模型将公开于https://anticdimi.github.io/lexis。
原文摘要 · Abstract (English)
Reconstructing 3D Human-Object Interaction from an RGB image is essential for perceptive systems. Yet, this remains challenging as it requires capturing the subtle physical coupling between the body and objects. While current methods rely on sparse, binary contact cues, these fail to model the continuous proximity and dense spatial relationships that characterize natural interactions. We address this limitation via InterFields, a representation that encodes dense, continuous proximity across the entire body and object surfaces. However, inferring these fields from single images is inherently ill-posed. To tackle this, our intuition is that interaction patterns are characteristically structured by the action and object geometry. We capture this structure in LEXIS, a novel discrete manifold of interaction signatures learned via a VQ-VAE. We then develop LEXIS-Flow, a diffusion framework that leverages LEXIS signatures to estimate human and object meshes alongside their InterFields. Notably, these InterFields help in a guided refinement that ensures physically-plausible, proximity-aware reconstructions without requiring post-hoc optimization. Evaluation on Open3DHOI and BEHAVE shows that LEXIS-Flow significantly outperforms existing SotA baselines in reconstruction, contact, and proximity quality. Our approach not only improves generalization but also yields reconstructions perceived as more realistic, moving us closer to holistic 3D scene understanding. Code & models will be public at https://anticdimi.github.io/lexis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。