分析视觉注意力头,揭示语言模型如何理解图像
Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach
- 通过分析4类模型、4种规模的注意力头,发现专用于视觉内容的特殊注意力机制
- 视觉注意力头的权重分布与视觉标记集中度高度相关
- 为构建多模态智能系统提供理论支持,适合关注视觉理解的研究者
多模态大语言模型(MLLMs)在视觉理解方面取得了显著进展。这一突破引发了一个关键问题:仅在语言数据上训练的语言模型如何有效解析和处理视觉内容?本文系统研究了4个模型家族和4种模型规模,揭示了一类专注于视觉内容的特殊注意力头。分析表明,这些注意力头的行为、注意力权重分布与其对输入中视觉标记的集中程度存在强相关性。该发现深化了我们对语言模型适应多模态任务的理解,展示了其在连接文本与视觉理解方面的潜力。本研究为开发能够处理多种模态的AI系统奠定了基础。
原文摘要 · Abstract (English)
Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated remarkable progress in visual understanding. This impressive leap raises a compelling question: how can language models, initially trained solely on linguistic data, effectively interpret and process visual content? This paper aims to address this question with systematic investigation across 4 model families and 4 model scales, uncovering a unique class of attention heads that focus specifically on visual content. Our analysis reveals a strong correlation between the behavior of these attention heads, the distribution of attention weights, and their concentration on visual tokens within the input. These findings enhance our understanding of how LLMs adapt to multimodal tasks, demonstrating their potential to bridge the gap between textual and visual understanding. This work paves the way for the development of AI systems capable of engaging with diverse modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。