提出TAO方法,让热成像定位系统在陌生环境不乱认路,还能快速找回迷路机器人。
Trajectory-Anchor Optimization for Overconfident Thermal Visual Place Recognition: Zero-Leakage OOD Auditing and Kidnapped-Robot Recovery

- 用批量坐标对齐取代传统多假设追踪,大幅降低计算开销。
- 在5米内能识别出虚假匹配,在10米外可精准发现定位错误。
- 适合需要高可靠性的自动驾驶与无人机等实时导航场景。
基于基础模型的现代热成像视觉定位(TIR-VPR)前端在闭集检索中表现优异,但在分布外(OOD)或未建图环境下会生成高度可信却错误的闭环候选,且相似度分数不下降。经典多假设追踪(MHT)后端虽可缓解歧义,但其指数级计算复杂度无法满足实时机器人需求。为此,本文提出轨迹锚点优化(TAO),将多视角时间验证压缩为批量SE(2)普鲁克斯特斯对齐问题。通过张量级向量化和单次调用的批量SVD,TAO避免了MHT的动态树扩展,保证每帧执行时间严格控制在O(KN)。在严格的零漏检评估下,被动几何后端因局部视觉模糊,无法在微尺度(<5米)区分度量误差与一致幻觉;而TAO作为高效容错机制,在宏观尺度起作用:5米内幻觉具有一致几何特征,欺骗刚性对齐;但超过此阈值后,K=100个假设在全局地图上分散,导致滑动窗口(N=20)内联合优化残差急剧上升,从而在10米范围内建立显著的宏观收敛区域,可靠分离灾难性拓扑断裂并抑制关键误接受。
原文摘要 · Abstract (English)
Modern thermal visual place recognition (TIR-VPR) frontends based on foundation models achieve remarkable closed-set retrieval but suffer from an overconfident forced-matching failure mode. Under out-of-distribution (OOD) or unmapped conditions, they generate highly plausible yet false loop candidates without a drop in similarity scores. While classical multi-hypothesis tracking (MHT) backends can mitigate these ambiguities by maintaining divergent trajectory beliefs, their exponential computational overhead violates real-time robotic constraints. To bridge this gap, we present Trajectory-Anchor Optimization (TAO). To counter the combinatorial challenge of evaluating parallel hypotheses (e.g., K=100), TAO compresses multi-view temporal verification into a batched SE(2) Procrustes alignment problem. By leveraging tensor-level vectorization and single-invocation batched SVD, this formulation bypasses the dynamic tree expansion of MHT, guaranteeing a strictly bounded per-frame execution loop of O(KN). Under a strict zero-leakage evaluation protocol, we show that while a passive geometric backend cannot mathematically separate metric localization errors from coherent hallucinations at a micro-scale (<5m) due to local visual ambiguities, TAO serves as an efficient fail-safe filter at a macro-scale. Within a 5m radius, hallucinations often possess a locally consistent geometry that deceives rigid alignment. However, beyond this threshold, the K=100 disparate hypotheses disperse spatially across the global map. This dispersion breaks the rigid temporal co-visibility constraint within the sliding window (N=20), causing the joint optimization residual to escalate sharply. Consequently, TAO establishes a distinct macroscopic convergence basin (10m) where multi-view geometric consistency reliably isolates catastrophic topological breaks and suppresses critical false acceptances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。