不训练就能提升大模型对图像小细节的识别能力
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
- 利用注意力与梯度图定位视觉关键区域
- 在7个评测集上显著提升小细节识别准确率
- 无需训练,适合部署于现有大模型系统
多模态大语言模型(MLLMs)近年来在视觉识别任务中进展迅速。然而,其在处理图像中小尺寸视觉对象时表现显著下降。我们发现该现象具有因果性,并通过干预实验验证。进一步分析显示,尽管回答错误,模型仍能准确定位目标位置。基于此,我们提出无需训练的视觉干预方法,利用模型内部的注意力与梯度信息增强对微小细节的感知。在两个主流MLLM和七个视觉问答基准上评估,结果表明该方法显著提升准确率。研究揭示了MLLM在细粒度视觉任务中的风险,同时指出利用模型内状态进行干预是有效缓解路径。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have experienced rapid progress in visual recognition tasks in recent years. Given their potential integration into many critical applications, it is important to understand the limitations of their visual perception. In this work, we study whether MLLMs can perceive small visual details as effectively as large ones when answering questions about images. We observe that their performance is very sensitive to the size of the visual subject of the question, and further show that this effect is in fact causal by conducting an intervention study. Next, we study the attention patterns of MLLMs when answering visual questions, and intriguingly find that they consistently know where to look, even when they provide the wrong answer. Based on these findings, we then propose training-free visual intervention methods that leverage the internal knowledge of any MLLM itself, in the form of attention and gradient maps, to enhance its perception of small visual details. We evaluate our proposed methods on two widely-used MLLMs and seven visual question answering benchmarks and show that they can significantly improve MLLMs' accuracy without requiring any training. Our results elucidate the risk of applying MLLMs to visual recognition tasks concerning small details and indicate that visual intervention using the model's internal state is a promising direction to mitigate this risk.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。