融合多模型互补特征,提升视觉定位在变化环境下的鲁棒性
DC-VLAQ: Query-Residual Aggregation for Robust Visual Place Recognition
- 用残差修正融合DINOv2与CLIP特征,保留语义互补性
- 提出查询残差聚合机制,稳定捕捉细粒度判别线索
- 在6个基准上表现领先,尤其适应光照、视角大变
视觉定位(VPR)的核心挑战在于学习在视角剧变、光照变化和严重域偏移下仍具判别力的全局表征。尽管视觉基础模型(VFMs)提供强局部特征,但现有方法多依赖单一模型,忽视不同VFMs间的互补信息。然而,利用这些互补信息会改变标记分布,影响现有基于查询的全局聚合方案的稳定性。为此,我们提出DC-VLAQ,一种以表征为中心的框架,整合互补VFMs融合与鲁棒全局聚合。首先,引入轻量级残差引导互补融合,在DINOv2特征空间中锚定表示,通过学习的残差校正注入来自CLIP的互补语义。其次,提出向量化局部聚合查询(VLAQ),通过可学习查询对局部标记的残差响应进行编码,实现更高稳定性和细粒度判别线索的保持。在Pitts30k、Tokyo24/7、MSLS、Nordland、SPED和AmsterTime等标准VPR基准上的大量实验表明,DC-VLAQ持续优于强基线,尤其在挑战性域偏移和长期外观变化下达到最先进性能。
原文摘要 · Abstract (English)
One of the central challenges in visual place recognition (VPR) is learning a robust global representation that remains discriminative under large viewpoint changes, illumination variations, and severe domain shifts. While visual foundation models (VFMs) provide strong local features, most existing methods rely on a single model, overlooking the complementary cues offered by different VFMs. However, exploiting such complementary information inevitably alters token distributions, which challenges the stability of existing query-based global aggregation schemes. To address these challenges, we propose DC-VLAQ, a representation-centric framework that integrates the fusion of complementary VFMs and robust global aggregation. Specifically, we first introduce a lightweight residual-guided complementary fusion that anchors representations in the DINOv2 feature space while injecting complementary semantics from CLIP through a learned residual correction. In addition, we propose the Vector of Local Aggregated Queries (VLAQ), a query--residual global aggregation scheme that encodes local tokens by their residual responses to learnable queries, resulting in improved stability and the preservation of fine-grained discriminative cues. Extensive experiments on standard VPR benchmarks, including Pitts30k, Tokyo24/7, MSLS, Nordland, SPED, and AmsterTime, demonstrate that DC-VLAQ consistently outperforms strong baselines and achieves state-of-the-art performance, particularly under challenging domain shifts and long-term appearance changes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。