arXiv:2601.22483cs.CV2026-01中稿 · 2026 IEEE Internat…被引 1

通过筛选注意力头提升视觉定位精度,让大模型更准回答细粒度问题。

Head-Aware Visual Cropping: Enhancing Fine-Grained VQA with Attention-Guided Subimage

  • 筛选具备真实视觉定位能力的注意力头,仅保留有效信号。
  • 在推理时进一步优化头的注意力分布,生成精准的裁剪指引图。
  • 无需训练,适配各类多模态大模型,提升细粒度问答准确率。

多模态大语言模型在视觉问答任务中表现强劲,但在细粒度推理上受限于低分辨率输入和噪声注意力聚合。本文提出无需训练的「头感知视觉裁剪」(HAVC)方法,通过基于OCR的诊断任务筛选出具有真实视觉定位能力的注意力头。推理时,利用空间熵增强空间集中性,通过梯度敏感性评估预测贡献,融合信号生成可靠的视觉裁剪引导图,精准定位任务相关区域,并指导子图像裁剪,随后将裁剪后的子图像与原图-问题对一同输入多模态大模型。在多个细粒度视觉问答基准测试中,HAVC持续优于现有裁剪策略,实现更精确的定位与更强的视觉对齐能力,为提升多模态大模型精度提供简单有效的方案。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) show strong performance in Visual Question Answering (VQA) but remain limited in fine-grained reasoning due to low-resolution inputs and noisy attention aggregation. We propose \textbf{Head Aware Visual Cropping (HAVC)}, a training-free method that improves visual grounding by leveraging a selectively refined subset of attention heads. HAVC first filters heads through an OCR-based diagnostic task, ensuring that only those with genuine grounding ability are retained. At inference, these heads are further refined using spatial entropy for stronger spatial concentration and gradient sensitivity for predictive contribution. The fused signals produce a reliable Visual Cropping Guidance Map, which highlights the most task-relevant region and guides the cropping of a subimage subsequently provided to the MLLM together with the image-question pair. Extensive experiments on multiple fine-grained VQA benchmarks demonstrate that HAVC consistently outperforms state-of-the-art cropping strategies, achieving more precise localization, stronger visual grounding, providing a simple yet effective strategy for enhancing precision in MLLMs.

视觉问答注意力机制视觉定位多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。