用几何方法提升视觉定位在极端变化下的鲁棒性
Beyond First-Order: Learning Riemannian Geometries for Invariant Visual Place Recognition
- 将场景特征建模为对称正定矩阵,利用黎曼几何捕捉二阶结构
- 零样本性能媲美有监督方法,微调后在非结构化环境达最新水平
- 适合需要高鲁棒性的自动驾驶与机器人定位任务
视觉定位(VPR)要求特征表示在剧烈环境和视角变化下仍保持稳定。现有聚合方法或依赖大量有监督训练,或采用一阶池化,难以在极端变化下保持结构关联性,且适应成本高。本文提出黎曼不变聚合(RIA),一种统一的几何框架,在对称正定(SPD)流形上显式建模二阶场景结构。通过将扰动视为可处理的合同变换,RIA利用几何感知的黎曼映射,将协方差描述子投影到线性欧氏空间,有效保留不变结构成分并抑制噪声。大量实验表明,RIA实现零样本性能媲美有监督方法,经简单微调后在非结构化环境中达到当前最优精度。源代码将公开。
原文摘要 · Abstract (English)
Visual Place Recognition (VPR) demands representations robust to drastic environmental and viewpoint shifts. Existing aggregation paradigms either depend on extensive supervised training or rely on first-order pooling, often struggling to preserve structural correlations under extreme shifts or incurring high adaptation costs. In this work, we propose Riemannian Invariant Aggregation (RIA), a unified geometric framework that explicitly models second-order scene structure on the Symmetric Positive Definite (SPD) manifold. By treating perturbations as tractable congruence transformations, RIA leverages geometry-aware Riemannian mappings to project covariance descriptors into a linearized Euclidean space, effectively preserving invariant structural components while suppressing noise. Extensive evaluations demonstrate that RIA achieves zero-shot performance comparable to supervised methods, and establishes state-of-the-art accuracy with simple fine-tuning, particularly in unstructured environments. The source code will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。