用2D扩散模型生成3D物体,提升视图一致性与细节质量
CDI3D: Cross-guided Dense-view Interpolation for 3D Reconstruction
- 引入稠密视图插值模块,增强多视角图像一致性
- 在多个基准上优于现有方法,纹理和几何精度显著提升
- 适合需要高质量3D重建的视觉生成与逆向建模研究者
从单张图像进行3D物体重建是计算机视觉中的基础任务,应用广泛。近年来,大型重建模型(LRMs)通过利用2D扩散模型生成的多视角图像提取3D内容展现出巨大潜力。然而,2D扩散模型常难以生成具有强多视角一致性的稠密图像,而LRMs在3D重建过程中会放大这些不一致性。为此,本文提出CDI3D,一种前馈式框架,实现高效、高质量的图像到3D生成与视图插值。我们设计了一个稠密视图插值(DVI)模块,在2D扩散模型生成的主视角之间合成中间视角,有效提升输入视图密度与一致性。同时,采用倾斜相机姿态轨迹捕获不同高度与视角的视图。随后,基于三平面策略从插值与原始视图中提取鲁棒特征,生成高质量3D网格。大量实验表明,该方法在多个基准上显著超越现有最先进方法,生成的3D内容在纹理保真度与几何准确性方面均有显著提升。
原文摘要 · Abstract (English)
3D object reconstruction from single-view image is a fundamental task in computer vision with wide-ranging applications. Recent advancements in Large Reconstruction Models (LRMs) have shown great promise in leveraging multi-view images generated by 2D diffusion models to extract 3D content. However, challenges remain as 2D diffusion models often struggle to produce dense images with strong multi-view consistency, and LRMs tend to amplify these inconsistencies during the 3D reconstruction process. Addressing these issues is critical for achieving high-quality and efficient 3D reconstruction. In this paper, we present CDI3D, a feed-forward framework designed for efficient, high-quality image-to-3D generation with view interpolation. To tackle the aforementioned challenges, we propose to integrate 2D diffusion-based view interpolation into the LRM pipeline to enhance the quality and consistency of the generated mesh. Specifically, our approach introduces a Dense View Interpolation (DVI) module, which synthesizes interpolated images between main views generated by the 2D diffusion model, effectively densifying the input views with better multi-view consistency. We also design a tilt camera pose trajectory to capture views with different elevations and perspectives. Subsequently, we employ a tri-plane-based mesh reconstruction strategy to extract robust tokens from these interpolated and original views, enabling the generation of high-quality 3D meshes with superior texture and geometry. Extensive experiments demonstrate that our method significantly outperforms previous state-of-the-art approaches across various benchmarks, producing 3D content with enhanced texture fidelity and geometric accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。