提出PoI过滤器,提升新视角合成的像素可靠性,增强3D场景定位精度。
PoI: A Filter to Extract Pixel of Interest from Novel Views for Scene Coordinate Regression

- 用3DGS渲染新视角,再通过扩散模型补全结构细节。
- 基于重投影误差逐像素筛选可信合成像素,减少错误信息干扰。
- 在7Scenes和Cambridge Landmarks上超越现有方法,适合高精度3D定位任务。
神经视图合成(NVS)技术如NeRF和3D高斯溅射(3DGS)已实现从新视角生成逼真图像,并被用于增强视觉定位的训练数据。然而,这些方法依赖已有几何与辐射信息,仅进行插值,无法生成未见的3D结构或恢复稀疏/极端视角下的缺失内容,导致渲染图像常出现模糊、结构失真或几何不完整。尽管相机位姿回归(CPR)可容忍此类瑕疵,但场景坐标回归(SCR)需要精确的像素级3D监督,因此性能严重受损。为此,本文提出PoI(Pixel-of-Interest),一个针对SCR定位的高效NVS增强框架。首先利用3DGS生成新视角,再通过单步扩散模型优化,合成更合理的结构细节;但即使经过扩散优化,仍存在不可靠像素。为此,我们设计一种基于重投影误差的渐进式像素级过滤策略,在训练中仅保留可信合成像素,抑制有害像素。在7Scenes和Cambridge Landmarks上的大量实验表明,该方法持续优于强基线模型,达到当前最优性能,且训练效率具有竞争力。结果揭示:对于SCR任务,新视角增强的收益不仅取决于生成真实性,更依赖对像素级可靠性的显式控制。
原文摘要 · Abstract (English)
Neural View Synthesis (NVS) techniques such as NeRF and 3D Gaussian Splatting (3DGS) have enabled photorealistic rendering from novel viewpoints and are increasingly used to augment training data for visual localization. However, these methods fundamentally rely on observed geometry and radiance; they interpolate existing information but cannot hallucinate unseen 3D structures or recover missing content under sparse or extreme viewpoints. As a result, rendered views often exhibit blur, structural distortion, or incomplete geometry. While such imperfections may be tolerated by Camera Pose Regression (CPR) methods, they severely degrade Scene Coordinate Regression (SCR), which requires accurate per-pixel 3D supervision. To address this limitation, we introduce PoI (Pixel-of-Interest), a framework that enables effective NVS augmentation for SCR-based localization. We first employ 3DGS to render novel views and leverage a single-step diffusion model to refine them, allowing the synthesis of structurally plausible details beyond purely geometry-driven interpolation. However, even diffusion-refined views may contain unreliable pixels. Therefore, we propose a progressive pixel-level filtering strategy based on reprojection error to selectively retain trustworthy synthetic pixels during training while suppressing harmful ones. Extensive experiments on 7Scenes and Cambridge Landmarks demonstrate that our method consistently improves localization accuracy over strong SCR baselines and achieves state-of-the-art performance with competitive training efficiency. Our results reveal that, for SCR, the benefit of novel view augmentation depends not only on generative realism but also on explicit control of pixel-level reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。