利用轨迹上下文提升弱视觉环境下的定位精度
TRAIL: Trajectory-Aware Visual Place Recognition against Unordered Databases

- 基于条件随机场融合视觉相似性与相机运动一致性
- 在无序数据库中定位准确率提升最高8.3个百分点
- 特别适合弱视觉场景,无需重新训练即可迁移
现代视觉位置识别方法在标准基准上表现优异,但在特征匮乏环境中仍显脆弱。由于将每个查询图像孤立处理,它们忽略了真实轨迹中的序列上下文。我们提出一个利用该上下文的任务:给定查询序列,在无序参考数据库中定位最后一张图像——与序列到序列方法不同,该数据库无需具备序列结构。我们提出TRAIL(轨迹感知图像定位),一种基于条件随机场(CRF)的原理性框架,结合学习得到的视觉相似性函数和相机运动一致性函数,随着每个查询图像的到来逐步优化候选参考的分布。TRAIL是任意预训练VPR主干网络上的轻量级后处理层,在主要基准上使最先进基线性能提升高达8.3个百分点,可在未见数据集上直接迁移,且在视觉线索稀缺时取得最大增益。
原文摘要 · Abstract (English)
Modern Visual Place Recognition (VPR) methods excel on standard benchmarks yet remain brittle in feature-poor environments. By treating each query image in isolation, they discard the sequential context in any real trajectory. We formalize a task that exploits this context: given a query sequence, localize the final image against an unordered reference database -- which, unlike sequence-to-sequence methods, requires no sequential structure in the database. We propose TRAIL (TRajectory-Aware Image Localization), a principled framework based on Conditional Random Fields (CRF) that combines learned functions for visual similarity and for camera-motion consistency, refining a distribution over candidate references as each query arrives. A lightweight post-processing layer atop any pre-trained VPR backbone, TRAIL improves a state-of-the-art baseline by up to 8.3 percentage points on our primary benchmark, transfers to unseen datasets without retraining, and delivers its largest gains where visual cues are scarce.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。