LaVPR用65万条语言描述提升视觉定位鲁棒性,让小模型也能抗环境变化。
LaVPR: Benchmarking Language and Vision for Place Recognition
- 融合语言与视觉信息增强定位鲁棒性
- 小模型加语言可媲美大视觉模型性能
- 适合应急救援等资源受限场景
视觉定位(VPR)在极端环境变化和感知歧义下表现不佳。现有系统无法仅通过语言描述实现‘盲定位’,而这对于应急响应等应用至关重要。为此,我们提出LaVPR,一个大规模基准,扩展了现有VPR数据集,包含超过65万条丰富的自然语言描述。利用LaVPR,我们研究了两种范式:多模态融合以提升鲁棒性,跨模态检索用于语言驱动定位。结果表明,语言描述在视觉退化条件下带来稳定提升,对小型骨干网络影响最显著。值得注意的是,加入语言后,紧凑模型的性能可媲美远大尺寸的纯视觉架构。对于跨模态检索,我们基于低秩适配(LoRA)和多相似性损失建立基线,显著优于标准对比方法。最终,LaVPR使新型定位系统成为可能——既具备应对真实世界随机性的鲁棒性,又适用于资源受限部署。数据集与代码已开源。
原文摘要 · Abstract (English)
Visual Place Recognition (VPR) often fails under extreme environmental changes and perceptual aliasing. Beyond these limitations, standard systems cannot perform 'blind' localization from verbal descriptions alone, a capability critical for applications such as emergency response. To address these challenges, we introduce LaVPR, a large-scale benchmark that extends existing VPR datasets with over 650,000 rich natural-language descriptions. Using LaVPR, we investigate two paradigms: Multi-Modal Fusion for enhanced robustness and Cross-Modal Retrieval for language-based localization. Our results show that language descriptions yield consistent gains in visually degraded conditions, with the most significant impact on smaller backbones. Notably, adding language allows compact models to rival the performance of much larger vision-only architectures. For cross-modal retrieval, we establish a baseline using Low-Rank Adaptation (LoRA) and Multi-Similarity loss, which substantially outperforms standard contrastive methods across vision-language models. Ultimately, LaVPR enables a new class of localization systems that are both resilient to real-world stochasticity and practical for resource-constrained deployment. Our dataset and code are available at https://github.com/oferidan1/LaVPR
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。