通过解耦视角与语义,实现稀疏视图下逼真3D重建。
VI3DRM:Towards meticulous 3D Reconstruction from Sparse Views via Photo-Realistic Novel View Synthesis

- 在解耦的3D潜在空间中进行扩散建模,分离光照、材质与结构
- 在GSO数据集上达成38.61的PSNR,超越现有最佳方法
- 适合需要高保真3D生成的视觉设计与数字孪生应用
近期基于单视图的3D重建方法如Zero-1-2-3取得了显著进展,但其对未见区域的预测高度依赖大规模预训练扩散模型的归纳偏置。尽管后续工作如DreamComposer尝试通过引入额外视图提升可控性,但由于原始潜在空间中存在光照、材质与结构等特征纠缠,结果仍不真实。为此,我们提出视觉各向同性3D重建模型(VI3DRM),一种基于扩散模型的稀疏视图3D重建方法,运行于身份一致且视角解耦的3D潜在空间中。该模型可有效解耦语义信息、颜色、材质属性和光照,从而生成近乎真实的图像。结合真实与合成图像,本方法能精确构建点云图,最终生成精细纹理的网格或点云。在GSO数据集的新型视图合成任务中,VI3DRM显著优于当前最优方法DreamComposer,达到38.61的PSNR、0.929的SSIM和0.027的LPIPS。代码将在发表后公开。
原文摘要 · Abstract (English)
Recently, methods like Zero-1-2-3 have focused on single-view based 3D reconstruction and have achieved remarkable success. However, their predictions for unseen areas heavily rely on the inductive bias of large-scale pretrained diffusion models. Although subsequent work, such as DreamComposer, attempts to make predictions more controllable by incorporating additional views, the results remain unrealistic due to feature entanglement in the vanilla latent space, including factors such as lighting, material, and structure. To address these issues, we introduce the Visual Isotropy 3D Reconstruction Model (VI3DRM), a diffusion-based sparse views 3D reconstruction model that operates within an ID consistent and perspective-disentangled 3D latent space. By facilitating the disentanglement of semantic information, color, material properties and lighting, VI3DRM is capable of generating highly realistic images that are indistinguishable from real photographs. By leveraging both real and synthesized images, our approach enables the accurate construction of pointmaps, ultimately producing finely textured meshes or point clouds. On the NVS task, tested on the GSO dataset, VI3DRM significantly outperforms state-of-the-art method DreamComposer, achieving a PSNR of 38.61, an SSIM of 0.929, and an LPIPS of 0.027. Code will be made available upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。