用法向图作桥梁,从图像生成高保真3D几何形状
Hi3DGen: High-fidelity 3D Geometry Generation from Images via Normal Bridging
- 通过噪声注入与双流训练分离图像频段,提升法向估计精度
- 利用法向正则化的潜空间扩散模型,显著增强3D形状细节还原能力
- 构建高质量3D数据集,适合需要高精度3D重建的研究者
随着从2D图像生成高保真3D模型的需求增长,现有方法仍因域差距和RGB图像固有的模糊性,在复现细粒度几何细节方面面临挑战。为此,我们提出Hi3DGen,一种通过法向图桥接实现高保真3D几何生成的新框架。该框架包含三个核心组件:(1) 图像到法向估计器,通过噪声注入与双流训练解耦低-高频图像模式,实现可泛化、稳定且锐利的估计;(2) 法向到几何学习方法,采用法向正则化的潜空间扩散学习,提升3D几何生成保真度;(3) 3D数据合成流水线,构建高质量训练数据集。大量实验表明,本框架在生成丰富几何细节方面表现优异,优于当前最优方法。本工作为基于图像的高保真3D几何生成提供了新思路,即利用法向图作为中间表征。
原文摘要 · Abstract (English)
With the growing demand for high-fidelity 3D models from 2D images, existing methods still face significant challenges in accurately reproducing fine-grained geometric details due to limitations in domain gaps and inherent ambiguities in RGB images. To address these issues, we propose Hi3DGen, a novel framework for generating high-fidelity 3D geometry from images via normal bridging. Hi3DGen consists of three key components: (1) an image-to-normal estimator that decouples the low-high frequency image pattern with noise injection and dual-stream training to achieve generalizable, stable, and sharp estimation; (2) a normal-to-geometry learning approach that uses normal-regularized latent diffusion learning to enhance 3D geometry generation fidelity; and (3) a 3D data synthesis pipeline that constructs a high-quality dataset to support training. Extensive experiments demonstrate the effectiveness and superiority of our framework in generating rich geometric details, outperforming state-of-the-art methods in terms of fidelity. Our work provides a new direction for high-fidelity 3D geometry generation from images by leveraging normal maps as an intermediate representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。