arXiv:2605.12399cs.CV2026-05International Conf…

用几何信息修复稀疏视角下的3D重建错误

GeoQuery: Geometry-Query Diffusion for Sparse-View Reconstruction

论文配图:GeoQuery: Geometry-Query Diffusion for Sparse-View Reconstruction
图 1 · 摘自论文原文
  • 通过几何对齐的查询机制替代损坏的渲染特征
  • 在局部窗口内聚合跨视图特征,减少错误匹配
  • 可无缝接入现有扩散模型,适用于极端稀疏视角

3D高斯点云(3DGS)已成为3D重建与新视角合成的主流方法,但在稀疏视角条件下仍易产生严重伪影。现有方法虽尝试用图像扩散模型修正渲染伪影,但通常依赖多视角自注意力从参考图像中检索信息,当3DGS生成的新视角严重失真时,受损的查询特征会导致错误的跨视图匹配,进而引发不一致的修复结果。为此,我们提出GeoQuery,一种基于几何引导的扩散框架,通过新颖的几何引导跨视图注意力(GCA)机制,将生成先验与显式几何线索结合。首先,利用预测深度图和相机位姿构建几何诱导对应场,采样参考特征形成几何对齐的代理查询,替代受损的渲染特征;其次,设计新的跨视图特征聚合流程,将跨视图注意力限制在每个代理查询附近的局部窗口内,有效获取有用特征并抑制虚假匹配。GeoQuery可无缝集成至现有基于扩散的管线中,在极端视角稀疏条件下仍能实现鲁棒重建。在稀疏视角新视角合成与渲染伪影消除任务上的大量实验验证了该方法的有效性。

原文摘要 · Abstract (English)

3D Gaussian Splatting (3DGS) has emerged as a prominent paradigm for 3D reconstruction and novel view synthesis. However, it remains vulnerable to severe artifacts when trained under sparse-view constraints. While recent methods attempt to rectify artifacts in rendered views using image diffusion models, they typically rely on multi-view self-attention to retrieve information from reference images. We observe that this mechanism often fails when the rendered novel views output by 3DGS are heavily corrupted: damaged query features lead to erroneous cross-view retrieval, resulting in inconsistent rendering refinement. To address this, we propose GeoQuery, a geometry-guided diffusion framework that integrates generative priors with explicit geometric cues via a novel Geometry-guided Cross-view Attention (GCA) mechanism. First, by leveraging predicted depth maps and camera poses, we construct a geometry-induced correspondence field to sample reference features, forming a geometry-aligned proxy query that replaces the corrupted rendering features. Furthermore, we design a new cross-view feature aggregation pipeline, in which we restrict the cross-view attention to a local window around each proxy query to effectively retrieve useful features while suppressing spurious matches. GeoQuery can be seamlessly integrated into existing diffusion-based pipelines, enabling robust reconstruction even under extreme view sparsity. Extensive experiments on sparse-view novel view synthesis and rendering artifact removal demonstrate the effectiveness of our approach.

3D重建扩散模型几何先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。