单图生成逼真3D人像,解决面部畸变与姿势不自然问题
PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit Remeshing
- 跨尺度扩散建模全身与面部特征联合分布
- 在CAPE和THuman2.1上实现高保真几何与纹理重建
- 适合虚拟试衣、影视特效等需要高质量3D人像的场景
详细且逼真的3D人体建模对诸多应用至关重要,但仅凭单张RGB图像完成全身重建仍具挑战,主要源于问题本身的病态性以及衣物拓扑复杂、自遮挡严重。本文提出PSHuman,一种利用多视角扩散模型先验显式重建人体网格的新框架。发现直接将多视角扩散应用于单视图人像会导致严重几何畸变,尤其在生成面部时。为此,提出跨尺度扩散模型,联合建模全局全身形状与局部面部特征的概率分布,实现无几何畸变、身份一致的新型视角生成。此外,为增强不同姿态下跨视角的身体形状一致性,将生成模型条件化于参数化人体模型如SMPL-X,利用其人体先验防止生成不符合解剖结构的异常视角。基于生成的多视角法向量与颜色图,采用以SMPL-X初始化的显式人体雕刻方法,高效恢复真实纹理人体网格。在CAPE与THuman2.1数据集上的大量实验与定量评估表明,PSHuman在几何细节、纹理保真度及泛化能力方面均表现优越。
原文摘要 · Abstract (English)
Detailed and photorealistic 3D human modeling is essential for various applications and has seen tremendous progress. However, full-body reconstruction from a monocular RGB image remains challenging due to the ill-posed nature of the problem and sophisticated clothing topology with self-occlusions. In this paper, we propose PSHuman, a novel framework that explicitly reconstructs human meshes utilizing priors from the multiview diffusion model. It is found that directly applying multiview diffusion on single-view human images leads to severe geometric distortions, especially on generated faces. To address it, we propose a cross-scale diffusion that models the joint probability distribution of global full-body shape and local facial characteristics, enabling detailed and identity-preserved novel-view generation without any geometric distortion. Moreover, to enhance cross-view body shape consistency of varied human poses, we condition the generative model on parametric models like SMPL-X, which provide body priors and prevent unnatural views inconsistent with human anatomy. Leveraging the generated multi-view normal and color images, we present SMPLX-initialized explicit human carving to recover realistic textured human meshes efficiently. Extensive experimental results and quantitative evaluations on CAPE and THuman2.1 datasets demonstrate PSHumans superiority in geometry details, texture fidelity, and generalization capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。