单图生成3D人体,通过多视角扩散模型提升重建质量
Intrinsic Geometry-Appearance Consistency Optimization for Sparse-View Gaussian Splatting
- 用多视角扩散模型从单图生成多视角图像,融入三维结构先验
- 结合深度感知的面部优化模块,减少人脸失真,提升整体精度
- 适合需要高质量3D人体重建的研究者或数字人应用开发
从单张图像进行3D人体重建是一项挑战性任务,以往研究多集中于此。近期方法尝试利用扩散模型引导,通过分数蒸馏采样(SDS)优化3D表示或生成后视图以辅助重建。然而这些方法常产生不良伪影(如人体结构扁平化、过度平滑),且在真实场景中泛化能力差。本文提出MVD-HuGaS,实现仅凭单张参考图像即可生成自由视角的3D人体渲染。首先,基于一个在高质量3D人体数据集上微调的增强型多视角扩散模型,从单图生成多视角图像,融合3D几何先验与人体结构先验。为从稀疏生成的多视角图像中准确推断相机位姿,引入对齐模块,联合优化3D高斯与相机位姿。此外,提出基于深度的面部畸变缓解模块,精修生成的人脸区域,提升整体重建保真度。最终,结合优化后的多视角图像及其精确相机位姿,对目标人体的3D高斯进行优化,实现高保真自由视角渲染。在Thuman2.0和2K2K数据集上的大量实验表明,MVD-HuGaS在单视角3D人体重建任务上达到当前最优性能。
原文摘要 · Abstract (English)
3D human reconstruction from a single image is a challenging problem and has been exclusively studied in the literature. Recently, some methods have resorted to diffusion models for guidance, optimizing a 3D representation via Score Distillation Sampling(SDS) or generating a back-view image for facilitating reconstruction. However, these methods tend to produce unsatisfactory artifacts (\textit{e.g.} flattened human structure or over-smoothing results caused by inconsistent priors from multiple views) and struggle with real-world generalization in the wild. In this work, we present \emph{MVD-HuGaS}, enabling free-view 3D human rendering from a single image via a multi-view human diffusion model. We first generate multi-view images from the single reference image with an enhanced multi-view diffusion model, which is well fine-tuned on high-quality 3D human datasets to incorporate 3D geometry priors and human structure priors. To infer accurate camera poses from the sparse generated multi-view images for reconstruction, an alignment module is introduced to facilitate joint optimization of 3D Gaussians and camera poses. Furthermore, we propose a depth-based Facial Distortion Mitigation module to refine the generated facial regions, thereby improving the overall fidelity of the reconstruction. Finally, leveraging the refined multi-view images, along with their accurate camera poses, MVD-HuGaS optimizes the 3D Gaussians of the target human for high-fidelity free-view renderings. Extensive experiments on Thuman2.0 and 2K2K datasets show that the proposed MVD-HuGaS achieves state-of-the-art performance on single-view 3D human rendering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。