arXiv:2503.12929cs.CV2025-03ICCV被引 17

通过逐步预测下一视角,实现单图生成一致的3D物体。

AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction

  • 先生成近输入视角的视图,再逐步合成远视角。
  • 在多个数据集上提升视图一致性,生成高质量3D资产。
  • 适合需要高保真3D重建的工业设计与虚拟现实场景。

新颖视图合成(NVS)是图像到3D生成的核心技术。然而,现有方法在相机姿态差异较大时,难以保持生成视图与输入视图的一致性,导致3D几何和纹理质量差。我们观察发现,目标视图越接近输入视图,其保真度越高。基于此,我们提出AR-1-to-3,一种基于扩散模型的新型逐视图预测范式:首先生成靠近输入视角的视图,将其作为上下文信息,逐步合成更远视角。为编码生成视图序列作为局部和全局条件,我们设计了堆叠局部特征编码(Stacked-LE)和基于LSTM的全局特征编码(LSTM-GE)。大量实验表明,该方法显著提升了生成视图与输入视图的一致性,生成了高保真3D资产。

原文摘要 · Abstract (English)

Novel view synthesis (NVS) is a cornerstone for image-to-3d creation. However, existing works still struggle to maintain consistency between the generated views and the input views, especially when there is a significant camera pose difference, leading to poor-quality 3D geometries and textures. We attribute this issue to their treatment of all target views with equal priority according to our empirical observation that the target views closer to the input views exhibit higher fidelity. With this inspiration, we propose AR-1-to-3, a novel next-view prediction paradigm based on diffusion models that first generates views close to the input views, which are then utilized as contextual information to progressively synthesize farther views. To encode the generated view subsequences as local and global conditions for the next-view prediction, we accordingly develop a stacked local feature encoding strategy (Stacked-LE) and an LSTM-based global feature encoding strategy (LSTM-GE). Extensive experiments demonstrate that our method significantly improves the consistency between the generated views and the input views, producing high-fidelity 3D assets.

3D生成扩散模型视图一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。