提出SPAR方法,让多模态模型更关注真实视觉信息。
Disentangling Semantic Attention from Structural Bias in the Attention Manifold

- 将注意力偏差视为结构化噪声,统一处理非语义视觉标记
- 净化干扰信号并重分配注意力,显著减少幻觉现象
- 无需训练、开销小,适合各类多模态模型快速部署
多模态大语言模型(MLLMs)中注意力机制虽表现优异,却常因对某些语义无用的视觉标记过度关注而产生偏差,称为‘视觉注意力黑洞’。现有干预方法多孤立处理这些标记,效率低下。本文将此现象视为一种广泛存在的结构化语言偏见,导致视觉语义信号被稀释,引发多模态幻觉。为此,提出无需训练、即插即用的Saliency-guided Purification and Adaptive Redistribution(SPAR)方法:先净化结构噪声,再将释放出的注意力预算重新分配至最具信息量的视觉区域。在多种幻觉评测基准上验证表明,SPAR能有效恢复真实的视觉锚定,且计算开销极低。
原文摘要 · Abstract (English)
The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disproportionate attention toward certain semantically uninformative visual tokens, a phenomenon termed "register" or "Visual Attention Sinks." While existing inference intervention methods attempt to identify these sink tokens and redistribute their attention weights, such approaches typically treat these tokens in isolation and suffer from computational inefficiency. Instead, we reframe this phenomenon as a generalized textual bias exerted over visual features that extends beyond isolated sink tokens. From this perspective, a pervasive structural bias leads to the dilution of the semantic visual signal, precipitating multimodal hallucinations as the model prioritizes linguistic priors over valid visual evidence. To address this limitation, we introduce Saliency-guided Purification and Adaptive Redistribution (SPAR), a training-free, plug-and-play intervention. SPAR mitigates this generalized textual bias by purifying structural noise and subsequently redistributing the reclaimed attention budget to the most informative visual regions. Comprehensive evaluations across a diverse spectrum of hallucination benchmarks demonstrate that SPAR effectively restores authentic visual grounding with negligible computational overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。