arXiv:2504.03536cs.CV2025-04被引 13

用高斯点云重建单图真人,效果更真实、肢体更完整。

HumanDreamer-X: Photorealistic Single-image Human Avatars Reconstruction via Gaussian Restoration

  • 用3D高斯泼溅做基础,统一生成与重建流程。
  • 修复高斯渲染结果,提升细节清晰度,PSNR达25.62 dB。
  • 解决多视角注意力缺陷,适合真实场景和多种模型使用。

单图像人体重建对数字人建模至关重要,但仍是极具挑战的任务。现有方法依赖生成模型合成多视角图像以进行后续3D重建与动画,但直接从单张人体图像生成多视图易导致几何不一致,出现肢体断裂或模糊等问题。为此,我们提出 extbf{HumanDreamer-X},一个将多视角人体生成与重建整合到统一流程的新型框架,显著提升了重建3D模型的几何一致性与视觉保真度。该框架采用3D高斯泼溅(3DGS)作为显式3D表示,提供初始几何与外观先验。在此基础上,训练 extbf{HumanFixer} 修复3DGS渲染结果,确保照片级真实感。此外,我们深入分析了多视角人体生成中注意力机制的内在挑战,提出一种注意力调制策略,有效增强了多视角间几何细节的身份一致性。实验表明,本方法在生成与重建的PSNR指标上分别提升16.45%与12.65%,最高达到25.62 dB,同时在真实场景数据上具备良好泛化能力,并可适配多种人体重建主干模型。

原文摘要 · Abstract (English)

Single-image human reconstruction is vital for digital human modeling applications but remains an extremely challenging task. Current approaches rely on generative models to synthesize multi-view images for subsequent 3D reconstruction and animation. However, directly generating multiple views from a single human image suffers from geometric inconsistencies, resulting in issues like fragmented or blurred limbs in the reconstructed models. To tackle these limitations, we introduce \textbf{HumanDreamer-X}, a novel framework that integrates multi-view human generation and reconstruction into a unified pipeline, which significantly enhances the geometric consistency and visual fidelity of the reconstructed 3D models. In this framework, 3D Gaussian Splatting serves as an explicit 3D representation to provide initial geometry and appearance priority. Building upon this foundation, \textbf{HumanFixer} is trained to restore 3DGS renderings, which guarantee photorealistic results. Furthermore, we delve into the inherent challenges associated with attention mechanisms in multi-view human generation, and propose an attention modulation strategy that effectively enhances geometric details identity consistency across multi-view. Experimental results demonstrate that our approach markedly improves generation and reconstruction PSNR quality metrics by 16.45% and 12.65%, respectively, achieving a PSNR of up to 25.62 dB, while also showing generalization capabilities on in-the-wild data and applicability to various human reconstruction backbone models.

3D重建高斯泼溅图像生成数字人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。