arXiv:2512.17226cs.CV2025-12中稿 · IEEE/CVF Winter Co…被引 1

用几何与视觉双重约束提升场景坐标回归的鲁棒性

Robust Scene Coordinate Regression via Geometrically-Consistent Global Descriptors

  • 融合几何结构与视觉相似性学习全局描述符
  • 在大规模基准上实现显著定位精度提升
  • 无需人工标签,适合跨环境部署

基于学习的视觉定位方法常使用全局描述符来区分视觉相似的位置,但现有方法多仅依赖几何线索(如共视图),导致判别力不足,在几何约束噪声下鲁棒性差。本文提出一个聚合模块,学习同时符合几何结构与视觉相似性的全局描述符,确保图像在描述符空间中接近仅当它们在视觉上相似且空间相连。该设计纠正了由不可靠重叠分数引发的错误关联。采用仅基于重叠分数的批处理挖掘策略和改进的对比损失,方法无需人工地点标签即可训练,并在多样环境中具有良好泛化能力。在多个挑战性基准测试中,该方法在大规模环境下实现了显著的定位性能提升,同时保持计算与内存效率。代码已开源:https://github.com/sontung/robust_scr。

原文摘要 · Abstract (English)

Recent learning-based visual localization methods use global descriptors to disambiguate visually similar places, but existing approaches often derive these descriptors from geometric cues alone (e.g., covisibility graphs), limiting their discriminative power and reducing robustness in the presence of noisy geometric constraints. We propose an aggregator module that learns global descriptors consistent with both geometrical structure and visual similarity, ensuring that images are close in descriptor space only when they are visually similar and spatially connected. This corrects erroneous associations caused by unreliable overlap scores. Using a batch-mining strategy based solely on the overlap scores and a modified contrastive loss, our method trains without manual place labels and generalizes across diverse environments. Experiments on challenging benchmarks show substantial localization gains in large-scale environments while preserving computational and memory efficiency. Code is available at https://github.com/sontung/robust_scr.

视觉定位全局描述符几何一致性无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。