arXiv:2510.07723cs.CV2025-10NeurIPS被引 8

融合2D多视角与3D原生生成模型,实现单图高保真人体重建。

SyncHuman: Synchronizing 2D and 3D Generative Models for Single-view Human Reconstruction

  • 用2D多视角与3D原生模型协同生成,互补细节与结构一致性
  • 提出像素对齐的2D-3D同步注意力机制,实现几何对齐的3D形状与多视图图像
  • 通过特征注入将2D细节映射到3D网格,提升重建精度与真实感

从单张图像实现逼真全身人体3D重建是影视与游戏应用中的关键挑战,因固有的歧义性和严重自遮挡而困难重重。现有方法依赖SMPL估计和SMPL条件化图像生成模型来推测新视角,但受限于不准确的SMPL先验,难以处理复杂姿态与精细细节。本文提出SyncHuman,首次结合2D多视角生成模型与3D原生生成模型,实现即使在挑战性姿态下仍高质量的着装人体网格重建。2D多视角模型擅长捕捉精细2D细节但结构一致性差,3D原生模型生成粗略但结构一致的3D形状。通过整合两者优势,我们首先联合微调两模型,采用像素对齐的2D-3D同步注意力机制,生成几何对齐的3D形状与多视图图像。为进一步提升细节,引入特征注入机制,将2D多视角图像中的精细特征迁移至对齐的3D形状,实现精准高保真重建。大量实验表明,SyncHuman在几何精度与视觉保真度上均优于基线方法,展现了未来3D生成模型的可行方向。

原文摘要 · Abstract (English)

Photorealistic 3D full-body human reconstruction from a single image is a critical yet challenging task for applications in films and video games due to inherent ambiguities and severe self-occlusions. While recent approaches leverage SMPL estimation and SMPL-conditioned image generative models to hallucinate novel views, they suffer from inaccurate 3D priors estimated from SMPL meshes and have difficulty in handling difficult human poses and reconstructing fine details. In this paper, we propose SyncHuman, a novel framework that combines 2D multiview generative model and 3D native generative model for the first time, enabling high-quality clothed human mesh reconstruction from single-view images even under challenging human poses. Multiview generative model excels at capturing fine 2D details but struggles with structural consistency, whereas 3D native generative model generates coarse yet structurally consistent 3D shapes. By integrating the complementary strengths of these two approaches, we develop a more effective generation framework. Specifically, we first jointly fine-tune the multiview generative model and the 3D native generative model with proposed pixel-aligned 2D-3D synchronization attention to produce geometrically aligned 3D shapes and 2D multiview images. To further improve details, we introduce a feature injection mechanism that lifts fine details from 2D multiview images onto the aligned 3D shapes, enabling accurate and high-fidelity reconstruction. Extensive experiments demonstrate that SyncHuman achieves robust and photo-realistic 3D human reconstruction, even for images with challenging poses. Our method outperforms baseline methods in geometric accuracy and visual fidelity, demonstrating a promising direction for future 3D generation models.

3D生成人体重建多视角生成细节增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。