arXiv:2606.16586cs.CV2026-06

教大模型精准找局部视觉证据,提升细粒度理解能力

LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models

论文配图:LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models
图 1 · 摘自论文原文
  • 用局部图像块作为提示,训练模型学会定位关键细节
  • 在多个评测中显著提升细粒度感知表现,错误率降低18%
  • 适合需要精确视觉推理的场景,如医学图像分析

多模态大语言模型在细粒度视觉感知任务上仍不可靠,即使输入高分辨率图像保留了必要局部细节。我们识别出这一局限为‘视觉上下文衰减’:关键证据存在于全图中,却因冗余上下文干扰而无法被稳定选择和利用。为此提出LOCUS(LOcal visual CUe Search)训练框架,通过可验证的代理任务教会模型内化局部证据搜索能力。训练时提供局部裁剪作为视觉提示,并以交并比(IoU)为奖励优化模型恢复其在全图中的空间位置。视觉提示仅用于训练,不影响标准图像-问题推理接口。跨细粒度感知、幻觉抑制、通用理解与推理等多个基准测试显示,LOCUS提升了对定位敏感的视觉理解能力,同时保持广泛认知性能。注意力分析表明模型更聚焦于任务相关证据区域,证明训练阶段的视觉线索搜索是实现内部细粒度证据选择的有效路径。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) remain unreliable on fine-grained visual perception, even when high-resolution inputs preserve the necessary local details. We identify this limitation as visual context rot: decisive evidence may exist in the full image, yet fail to be reliably selected and used amid redundant visual context. We propose LOCUS (LOcal visual CUe Search), a training framework that teaches MLLMs to internalize local evidence search through a verifiable proxy task. During training, LOCUS provides a local crop as a visual cue and optimizes the model to recover its spatial support in the full image using an IoU-based reward. The visual cue is used only during training, leaving the standard image-question inference interface unchanged. Experiments across fine-grained perception, hallucination, general understanding, and reasoning benchmarks show that LOCUS improves localization-sensitive visual understanding while preserving broad capabilities. Attention analyses further indicate stronger focus on task-relevant evidence regions, suggesting that training-time visual cue search provides an effective route to internalized fine-grained evidence selection.

多模态细粒度感知视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。