仅用一张图实现高质量3D人体自由视角渲染
MVD-HuGaS: Human Gaussians from a Single Image via 3D Human Multi-view Diffusion Prior
- 用多视角扩散模型生成参考图像的多角度视图
- 通过姿态对齐优化3D高斯点云,重建更准确
- 专设面部修正模块,提升人脸细节真实度
单图3D人体重建是极具挑战性的任务。现有方法多依赖扩散模型引导,通过得分蒸馏采样(SDS)优化3D表示或生成后视图辅助重建,但常出现结构扁平、过度平滑等伪影,且在真实场景中泛化能力差。本文提出MVD-HuGaS,利用多视角人体扩散模型从单张参考图生成多视角图像,该模型在高质量3D人体数据集上微调,融合了3D几何与人体结构先验。为从稀疏生成视图中准确推断相机位姿,引入对齐模块联合优化3D高斯与相机参数。此外,设计基于深度的面部畸变缓解模块,精细化修复面部区域,提升整体重建保真度。最终,结合精修后的多视角图像与精确相机位姿,优化目标人体的3D高斯表示,实现高保真自由视角渲染。在Thuman2.0和2K2K数据集上的实验表明,MVD-HuGaS在单视图3D人体渲染任务中达到当前最优性能。
原文摘要 · Abstract (English)
3D human reconstruction from a single image is a challenging problem and has been exclusively studied in the literature. Recently, some methods have resorted to diffusion models for guidance, optimizing a 3D representation via Score Distillation Sampling(SDS) or generating one back-view image for facilitating reconstruction. However, these methods tend to produce unsatisfactory artifacts (\textit{e.g.} flattened human structure or over-smoothing results caused by inconsistent priors from multiple views) and struggle with real-world generalization in the wild. In this work, we present \emph{MVD-HuGaS}, enabling free-view 3D human rendering from a single image via a multi-view human diffusion model. We first generate multi-view images from the single reference image with an enhanced multi-view diffusion model, which is well fine-tuned on high-quality 3D human datasets to incorporate 3D geometry priors and human structure priors. To infer accurate camera poses from the sparse generated multi-view images for reconstruction, an alignment module is introduced to facilitate joint optimization of 3D Gaussians and camera poses. Furthermore, we propose a depth-based Facial Distortion Mitigation module to refine the generated facial regions, thereby improving the overall fidelity of the reconstruction.Finally, leveraging the refined multi-view images, along with their accurate camera poses, MVD-HuGaS optimizes the 3D Gaussians of the target human for high-fidelity free-view renderings. Extensive experiments on Thuman2.0 and 2K2K datasets show that the proposed MVD-HuGaS achieves state-of-the-art performance on single-view 3D human rendering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。