从单张图重建多人交互的高保真3D人体模型
Human Interaction-Aware 3D Reconstruction from a Single Image

- 将图像转为正交空间,联合建模个体与群体上下文
- 生成完整多视角法向与图像,解决遮挡与重叠问题
- 引入物理交互先验,确保人与人接触真实可信
从单张图像重建带纹理的3D人体模型是AR/VR和数字人应用的基础。然而,现有方法多聚焦于单人场景,难以处理多人交互时的复杂情况,简单拼接个体重建常导致不合理的重叠、遮挡区域缺失几何结构以及交互扭曲等伪影。这凸显了对融合群体上下文与交互先验的方法的需求。本文提出一个整体性框架,显式建模群体与个体层级信息。首先将输入图像转换至规范正交空间以缓解透视畸变。核心组件Human Group-Instance Multi-View Diffusion(HUG-MVD)通过联合建模个体与群体上下文,生成完整的多视角法向与图像,有效解决遮挡与邻近关系问题。随后,Human Group-Instance Geometric Reconstruction(HUG-GR)模块利用显式的物理交互先验优化几何结构,保证物理合理性并精确建模人与人之间的接触。最后,多视角图像融合为高保真纹理。整体框架名为HUG3D。大量实验表明,HUG3D显著优于单人及现有多人方法,在仅输入单张图像的情况下即可生成物理合理、高保真的多人交互3D重建结果。
原文摘要 · Abstract (English)
Reconstructing textured 3D human models from a single image is fundamental for AR/VR and digital human applications. However, existing methods mostly focus on single individuals and thus fail in multi-human scenes, where naive composition of individual reconstructions often leads to artifacts such as unrealistic overlaps, missing geometry in occluded regions, and distorted interactions. These limitations highlight the need for approaches that incorporate group-level context and interaction priors. We introduce a holistic method that explicitly models both group- and instance-level information. To mitigate perspective-induced geometric distortions, we first transform the input into a canonical orthographic space. Our primary component, Human Group-Instance Multi-View Diffusion (HUG-MVD), then generates complete multi-view normals and images by jointly modeling individuals and group context to resolve occlusions and proximity. Subsequently, the Human Group-Instance Geometric Reconstruction (HUG-GR) module optimizes the geometry by leveraging explicit, physics-based interaction priors to enforce physical plausibility and accurately model inter-human contact. Finally, the multi-view images are fused into a high-fidelity texture. Together, these components form our complete framework, HUG3D. Extensive experiments show that HUG3D significantly outperforms both single-human and existing multi-human methods, producing physically plausible, high-fidelity 3D reconstructions of interacting people from a single image. Project page: https://jongheean11.github.io/HUG3D_project
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。