arXiv:2602.12003cs.CV2026-02

用外部视觉表征引导扩散模型,提升新视角生成的几何一致性与质量。

Projected Representation Conditioning for High-fidelity Novel View Synthesis

  • 通过投影模块将外部表征注入扩散过程,增强几何与语义对应关系。
  • 在标准基准上显著提升重建精度与补全质量,优于现有扩散方法。
  • 适用于稀疏、无姿态图像集,适合高保真新视角合成场景。

我们提出一种基于扩散模型的新视角合成框架,利用外部视觉表征作为条件,借助其几何与语义对应特性,提升生成新视角的几何一致性。首先,我们详细分析了外部视觉表征中空间注意力机制所涌现的对应能力。基于这些洞察,我们设计了名为ReNoV(representation-guided novel view synthesis)的表征引导新视角合成方法,通过专用的表征投影模块将外部表征注入扩散过程。实验表明,该设计在重建保真度和修补质量上均有显著提升,在标准基准上超越了现有的扩散基新视角方法,并可实现从稀疏、无姿态图像集合中的稳健合成。

原文摘要 · Abstract (English)

We propose a novel framework for diffusion-based novel view synthesis in which we leverage external representations as conditions, harnessing their geometric and semantic correspondence properties for enhanced geometric consistency in generated novel viewpoints. First, we provide a detailed analysis exploring the correspondence capabilities emergent in the spatial attention of external visual representations. Building from these insights, we propose a representation-guided novel view synthesis through dedicated representation projection modules that inject external representations into the diffusion process, a methodology named ReNoV, short for representation-guided novel view synthesis. Our experiments show that this design yields marked improvements in both reconstruction fidelity and inpainting quality, outperforming prior diffusion-based novel-view methods on standard benchmarks and enabling robust synthesis from sparse, unposed image collections.

新视角合成扩散模型几何一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。