不微调模型,用测试时缩放实现高效跨域视觉定位
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
- 测试时通过结构化提示引导多模态大模型直接评分
- 跨域识别性能提升,计算效率提高210倍
- 无需训练,适合实时应用和资源受限场景
视觉位置识别(VPR)已从手工设计特征演进到深度学习方法,但仍面临显著挑战。现有方法包括视觉基础模型(VFMs)和多模态大语言模型(MLLMs),虽增强语义理解,但微调后存在高计算开销与跨域迁移能力有限的问题。为此,我们提出一种新颖的零样本框架——测试时缩放(TTS),利用MLLM的视觉-语言对齐能力,通过基于引导的方法实现直接相似性评分。该方法摒弃两阶段处理,采用结构化提示生成长度可控的JSON输出。结合不确定性感知自一致性(UASC)的TTS框架可在无额外训练成本下实现实时适应,实现跨环境优越泛化能力。实验表明,该方法在跨域VPR性能上显著提升,计算效率最高提升210倍。
原文摘要 · Abstract (English)
Visual Place Recognition (VPR) has evolved from handcrafted descriptors to deep learning approaches, yet significant challenges remain. Current approaches, including Vision Foundation Models (VFMs) and Multimodal Large Language Models (MLLMs), enhance semantic understanding but suffer from high computational overhead and limited cross-domain transferability when fine-tuned. To address these limitations, we propose a novel zero-shot framework employing Test-Time Scaling (TTS) that leverages MLLMs' vision-language alignment capabilities through Guidance-based methods for direct similarity scoring. Our approach eliminates two-stage processing by employing structured prompts that generate length-controllable JSON outputs. The TTS framework with Uncertainty-Aware Self-Consistency (UASC) enables real-time adaptation without additional training costs, achieving superior generalization across diverse environments. Experimental results demonstrate significant improvements in cross-domain VPR performance with up to 210$\times$ computational efficiency gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。