arXiv:2509.25916cs.CVcs.CL2025-09被引 10

让视觉语言模型精准定位物体,突破细粒度感知瓶颈。

VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs

  • 将坐标生成转为特征检索,提升定位鲁棒性。
  • 在多个基准上达最优,物体定位准确率显著提升。
  • 可插拔集成,不损害模型原有理解能力,适合实用部署。

视觉语言模型(VLMs)在高层次场景理解上表现优异,但在需要精确定位的细粒度感知任务中表现不佳,根源在于语言主导架构难以生成精确数值坐标。本文提出VLM-FO1框架,将对象中心感知从脆弱的坐标生成问题重构为稳健的特征检索任务。该方法作为即插即用模块,可集成于任意预训练VLM。其采用双视觉编码器的混合细粒度区域编码器(HFRE),生成富含语义与空间细节的区域标记;基于标记的引用系统使大语言模型能无缝推理并锚定语言到特定视觉区域。实验表明,VLM-FO1在多样化基准上达到领先性能,展现出卓越的物体定位、区域生成理解与视觉区域推理能力。关键的是,其两阶段训练策略确保感知提升不损害基础模型的通用视觉理解能力。VLM-FO1建立了一种高效灵活的感知增强范式,弥合了高层推理与细粒度视觉锚定之间的鸿沟。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) excel at high-level scene understanding but falter on fine-grained perception tasks requiring precise localization. This failure stems from a fundamental mismatch, as generating exact numerical coordinates is a challenging task for language-centric architectures. In this paper, we introduce VLM-FO1, a novel framework that overcomes this limitation by reframing object-centric perception from a brittle coordinate generation problem into a robust feature retrieval task. Our method operates as a plug-and-play module that integrates with any pre-trained VLM. It leverages a Hybrid Fine-grained Region Encoder (HFRE), featuring a dual vision encoder, to generate powerful region tokens rich in both semantic and spatial detail. A token-based referencing system then enables the LLM to seamlessly reason about and ground language in these specific visual regions. Experiments show that VLM-FO1 achieves state-of-the-art performance across a diverse suite of benchmarks, demonstrating exceptional capabilities in object grounding, region generational understanding, and visual region reasoning. Crucially, our two-stage training strategy ensures that these perception gains are achieved without compromising the base model's general visual understanding capabilities. VLM-FO1 establishes an effective and flexible paradigm for building perception-aware VLMs, bridging the gap between high-level reasoning and fine-grained visual grounding.

视觉定位多模态模型增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。