通过随机注意力机制增强视觉表征稳定性,减少幻觉与噪声干扰。
Information-Regularized Attention for Visual-Centric Reasoning

- 引入信息正则化注意力(IRA),在中间层主动控制视觉信息注入量。
- 实验显示IRA使视觉嵌入曲率更平滑,抑制了各层注意力集中现象。
- 适合关注视觉-语言模型稳定性和鲁棒性提升的研究者。
视觉-语言模型(VLMs)已成为多模态学习的主流范式,但仍面临物体幻觉、视觉定位薄弱以及全参数指令微调后的灾难性遗忘等问题。我们指出,这些缺陷源于标准的下一个词预测目标对视觉表征学习缺乏显式控制。因此,视觉嵌入被动优化,容易引入冗余或虚假信号。为此,我们提出信息正则化注意力(IRA),一种随机注意力机制,可显式调节中间Transformer层中视觉信息的注入量。该局部重参数化将视觉表示的不确定性转化为独立于数据点的局部噪声。除了评估模型性能外,我们还量化了嵌入特性:IRA使曲率轨迹更平滑,并抑制所有层中的注意力汇聚现象,表明视觉信号转换更为稳定。结果表明,随机注意力不仅是正则化手段,更是生成架构中表征学习的关键因素,为构建更可靠的VLMs提供了新方向。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have become a paradigm for multimodal learning, yet remain unstable due to object hallucination, weak visual grounding, and catastrophic forgetting after full-parameter instruction tuning. We claim these failures result from a lack of explicit control over visual representation learning during the standard next-token prediction objective. As a result, visual embeddings thus become passively optimized and prone to injecting redundant or spurious signals. To counter this, we introduce Information-Regularized Attention (IRA), a stochastic attention mechanism that explicitly regulates the amount of visual information injected into the hidden states of intermediate transformer layers. This local reparameterization translates uncertainty about visual representations into local noise that is independent across data points. Beyond evaluating model performance, we also quantify embedding properties, where IRA produces smoother curvature trajectories and suppresses attention-sink across all layers, indicating a more stable transformation of the visual signal. Our results suggest that stochastic attention is not merely a regularizer but a key contributor to representation learning in a generative architecture, offering a new direction for building more reliable VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。