arXiv:2507.23569cs.CV2025-07CVPR被引 10

用3D高斯点云结合特征场,实现高精度且保护隐私的视觉定位。

Gaussian Splatting Feature Fields for Privacy-Preserving Visual Localization

  • 融合3D高斯点云与隐式特征场,构建场景表示。
  • 在真实数据集上达到顶尖定位性能,支持隐私保护模式。
  • 适合需要高精度定位且关注数据隐私的应用场景。

视觉定位旨在估计相机在已知环境中的位姿。本文利用基于3D高斯溅射(3DGS)的表征实现高精度且隐私保护的视觉定位。提出高斯溅射特征场(GSFF),结合显式几何模型(3DGS)与隐式特征场,利用3DGS的稠密几何信息和可微渲染算法,学习基于3D的鲁棒特征表示。特别地,通过对比框架将3D尺度感知特征场与2D特征编码器对齐于统一嵌入空间。进一步采用3D结构引导聚类过程,正则化表示学习,并无缝将特征转化为分割结果,可用于隐私保护的视觉定位。通过将查询图像的特征图或分割结果与从GSFF生成的渲染结果对齐,完成位姿精修。在多个真实世界数据集上的实验表明,该方法在隐私与非隐私场景下均达到当前最优性能。

原文摘要 · Abstract (English)

Visual localization is the task of estimating a camera pose in a known environment. In this paper, we utilize 3D Gaussian Splatting (3DGS)-based representations for accurate and privacy-preserving visual localization. We propose Gaussian Splatting Feature Fields (GSFFs), a scene representation for visual localization that combines an explicit geometry model (3DGS) with an implicit feature field. We leverage the dense geometric information and differentiable rasterization algorithm from 3DGS to learn robust feature representations grounded in 3D. In particular, we align a 3D scale-aware feature field and a 2D feature encoder in a common embedding space through a contrastive framework. Using a 3D structure-informed clustering procedure, we further regularize the representation learning and seamlessly convert the features to segmentations, which can be used for privacy-preserving visual localization. Pose refinement, which involves aligning either feature maps or segmentations from a query image with those rendered from the GSFFs scene representation, is used to achieve localization. The resulting privacy- and non-privacy-preserving localization pipelines, evaluated on multiple real-world datasets, show state-of-the-art performances.

视觉定位3D高斯隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。