arXiv:2608.28383cs.CVcs.CL2026-08

发现视觉模型注意力头有对象与背景分工,可量化并指导高效混合注意力设计。

Semantic Head Specialization Guides Hybrid ViT Attention for Multimodal LLMs

  • 通过分析注意力头角色分化,提出可量化的语义头专一性指标
  • 新设计的Ariadne注意力在22项任务上性能媲美全注意力,计算量降低6.5倍
  • 适合关注多模态大模型高效推理与注意力机制优化的研究者

混合注意力主导前沿大语言模型,但多模态大模型中的视觉变换器(ViTs)缺乏满意的混合设计,且尚无共识解释为何某些注意力模式更优。本文研究ViT注意力头,发现其分化为对象与背景专家角色,这一模式在全注意力下最为显著,称之为语义头专一性(SHS)。提出SHS-Index量化该专一性,证明其能区分全注意力与块窗口式ViTs,并与下游基准表现强相关。进一步识别出三个塑造SHS的结构因素——窗口交互、标记序列化和局部Softmax分配——并以此为设计原则,构建Ariadne Attention,可在22个图像与视频任务上达到全注意力性能,同时将注意力计算量降低6.5倍。研究成果确立头专一性为可度量属性,可用于诊断与设计多模态大模型尺度下的合理混合注意力。

原文摘要 · Abstract (English)

Hybrid attention dominates frontier LLMs, yet Vision Transformers (ViTs) in multimodal LLMs lack a satisfactory hybrid design, with no consensus on why certain attention patterns work better. To fill this gap, we study ViT attention heads and find they differentiate into object- and background-specialist roles, a pattern most pronounced under full attention; we call this Semantic Head Specialization (SHS). We propose SHS-Index to quantify this specialization, show that it distinguishes full-attention from chunk-window ViTs, and find that it strongly tracks downstream benchmark performance. We then identify three structural factors that shape SHS---window interaction, token serialization, and local softmax allocation---and use them as design principles for hybrid attention. Guided by these factors, we design Ariadne Attention, a hybrid that matches full attention on 22 image and video tasks at 6.5x less attention compute. Our findings establish head specialization as a measurable property for diagnosing and designing principled hybrid ViT attention at the multimodal-LLM scale.

视觉变换器注意力机制多模态高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。