用扩散模型提升稀疏视角下人体三维重建精度与泛化能力
DiHuR: Diffusion-Guided Generalizable Human Reconstruction
- 通过SMPL顶点令牌聚合多视角特征,引导隐式距离场预测
- 在三个数据集上实现优于现有方法的跨身份泛化性能
- 结合前馈模型与扩散先验,无需3D标注即可训练
我们提出DiHuR,一种基于扩散模型的通用人体三维重建与视角合成方法,仅需稀疏、重叠极少的多视角图像。现有通用人体辐射场虽擅长新视角合成,但三维重建效果不佳;而直接从稀疏视角优化隐式符号距离函数(SDF)因视图重叠不足常导致结果差。为此,我们引入与SMPL顶点关联的可学习令牌,聚合多视角特征并指导SDF预测。这些令牌在训练数据中学习到跨身份的通用先验,利用不同人体身份下SMPL顶点在图像中的语义区域投影一致性,实现对未见身份的有效知识迁移。考虑到SMPL难以捕捉衣物细节,我们引入扩散模型作为额外先验,补充复杂衣物几何缺失信息。该方法融合前馈模型先验与2D扩散先验,在无3D监督条件下仅需多视角图像训练。在THuman、ZJU-MoCap和HuMMan数据集上的实验表明,DiHuR在同数据集与跨数据集泛化设置下均显著优于现有方法。
原文摘要 · Abstract (English)
We introduce DiHuR, a novel Diffusion-guided model for generalizable Human 3D Reconstruction and view synthesis from sparse, minimally overlapping images. While existing generalizable human radiance fields excel at novel view synthesis, they often struggle with comprehensive 3D reconstruction. Similarly, directly optimizing implicit Signed Distance Function (SDF) fields from sparse-view images typically yields poor results due to limited overlap. To enhance 3D reconstruction quality, we propose using learnable tokens associated with SMPL vertices to aggregate sparse view features and then to guide SDF prediction. These tokens learn a generalizable prior across different identities in training datasets, leveraging the consistent projection of SMPL vertices onto similar semantic areas across various human identities. This consistency enables effective knowledge transfer to unseen identities during inference. Recognizing SMPL's limitations in capturing clothing details, we incorporate a diffusion model as an additional prior to fill in missing information, particularly for complex clothing geometries. Our method integrates two key priors in a coherent manner: the prior from generalizable feed-forward models and the 2D diffusion prior, and it requires only multi-view image training, without 3D supervision. DiHuR demonstrates superior performance in both within-dataset and cross-dataset generalization settings, as validated on THuman, ZJU-MoCap, and HuMMan datasets compared to existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。