从单张图片重建复杂人群3D高斯点云,解决遮挡与模糊难题。
CrowdGaussian: Reconstructing High-Fidelity 3D Gaussians for Human Crowd from a Single Image
- 基于自监督适配和自校准学习,直接生成多人体3D高斯表示。
- 在严重遮挡下仍能还原完整人体几何与外观,重建质量显著提升。
- 适合需要高保真人群3D重建的影视、虚拟现实等应用。
单视角3D人体重建近年来备受关注。尽管进展众多,以往研究主要针对清晰、近景的单人图像,对更常见的多人场景效果不佳。多人场景重建面临三大挑战:1)严重遮挡,2)图像模糊,3)人体姿态与外观多样。为此,我们提出CrowdGaussian,一个统一框架,可直接从单张图像重建多人体3D高斯点阵(3DGS)表示。为应对遮挡,设计自监督适配流程,使预训练大模型能在高度遮挡输入下重建出合理几何与外观的完整人体。此外,引入自校准学习(SCL),通过融合身份保持样本与干净/污染图像对,使单步扩散模型自适应优化粗略渲染至最优质量,并将结果回传以增强多人体3DGS表示。大量实验表明,CrowdGaussian能生成逼真且几何一致的多人场景3D重建。
原文摘要 · Abstract (English)
Single-view 3D human reconstruction has garnered significant attention in recent years. Despite numerous advancements, prior research has concentrated on reconstructing 3D models from clear, close-up images of individual subjects, often yielding subpar results in the more prevalent multi-person scenarios. Reconstructing 3D human crowd models is a highly intricate task, laden with challenges such as: 1) extensive occlusions, 2) low clarity, and 3) numerous and various appearances. To address this task, we propose CrowdGaussian, a unified framework that directly reconstructs multi-person 3D Gaussian Splatting (3DGS) representations from single-image inputs. To handle occlusions, we devise a self-supervised adaptation pipeline that enables the pretrained large human model to reconstruct complete 3D humans with plausible geometry and appearance from heavily occluded inputs. Furthermore, we introduce Self-Calibrated Learning (SCL). This training strategy enables single-step diffusion models to adaptively refine coarse renderings to optimal quality by blending identity-preserving samples with clean/corrupted image pairs. The outputs can be distilled back to enhance the quality of multi-person 3DGS representations. Extensive experiments demonstrate that CrowdGaussian generates photorealistic, geometrically coherent reconstructions of multi-person scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。