arXiv:2508.00945cs.CV2025-08

通过分层区域注意力对齐,提升视觉语言模型的跨模态一致性。

Optimizing Vision-Language Consistency via Cross-Layer Regional Attention Alignment

  • 引入分层-区块交叉注意力,细粒度捕捉区域与语义关联。
  • 在10个基准上超越基线,仅增355万参数即达顶尖性能。
  • 适合关注模型可解释性与跨模态对齐的研究者。

视觉语言模型在协调多种注意力机制进行跨模态嵌入学习时面临挑战,导致注意力错配和性能不佳。本文提出一致的分层区域对齐(CCRA),引入层-区块交叉注意力(LPWCA),通过联合加权区块和层级嵌入来捕捉细粒度的区域-语义关联;同时设计渐进式注意力融合(PAI),按顺序系统性协调LPWCA、层级和区块注意力机制。该渐进式设计确保从语义到区域层面的一致性,防止注意力漂移并最大化各注意力机制的优势。在10个不同视觉语言基准上的实验表明,使用CCRA增强的LLaVA-v1.5-7B模型达到最新性能,优于所有基线方法,仅增加355万额外参数,并通过更聚焦区域且语义对齐的注意力模式提升了可解释性。

原文摘要 · Abstract (English)

Vision Language Models (VLMs) face challenges in effectively coordinating diverse attention mechanisms for cross-modal embedding learning, leading to mismatched attention and suboptimal performance. We propose Consistent Cross-layer Regional Alignment (CCRA), which introduces Layer-Patch-wise Cross Attention (LPWCA) to capture fine-grained regional-semantic correlations by jointly weighting patch and layer-wise embedding, and Progressive Attention Integration (PAI) that systematically coordinates LPWCA, layer-wise, and patch-wise attention mechanisms in sequence. This progressive design ensures consistency from semantic to regional levels while preventing attention drift and maximizing individual attention benefits. Experimental results on ten diverse vision-language benchmarks demonstrate that our CCRA-enhanced LLaVA-v1.5-7B model achieves state-of-the-art performance, outperforming all baseline methods with only 3.55M additional parameters, while providing enhanced interpretability through more regionally focused and semantically aligned attention patterns.

视觉语言注意力对齐模型优化可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。