arXiv:2601.12430cs.CL2026-01ACL

发现视觉语言模型的'是'偏见源于系统注意力失衡,调节可有效抑制幻觉。

System-Mediated Attention Imbalances Make Vision-Language Models Say Yes

  • 提出系统中介注意力失衡新视角,指出系统权重冗余导致图像和文本关注不足。
  • 将系统注意力重分配后,'是'偏见显著降低,效果优于现有方法。
  • 适合关注模型幻觉、可解释性及多模态对齐的研究者阅读。

视觉语言模型(VLM)的幻觉常与输入模态间注意力分配不均相关:系统、图像和文本模态之间存在不平衡。然而,现有缓解策略多从图像中心视角出发,倾向于增强图像注意力,而忽视其他模态的作用。本研究提出更全面的系统中介视角,认为这种不平衡源于功能冗余的系统权重,其削弱了对图像和文本输入的关注。我们证明该框架为理解常见的'是'偏见——即模型无差别回答'是'——提供了有效的实证视角。因果地将注意力从系统模态重新分配至图像和文本输入,能显著抑制此偏见,且表现常优于现有方法。进一步证据表明,系统中介注意力失衡通过促使模型依赖粗粒度输入表示,引发默认响应倾向,这对某些任务有效,却在另一些任务中表现不佳。综合来看,这些发现确立了系统注意力在VLM幻觉中的关键作用,并揭示其作为缓解杠杆的巨大潜力。

原文摘要 · Abstract (English)

Vision-language model (VLM) hallucination is commonly linked to imbalanced allocation of attention across input modalities: system, image and text. However, existing mitigation strategies tend towards an image-centric interpretation of these imbalances, often prioritising increased image attention while giving less consideration to the roles of the other modalities. In this study, we evaluate a more holistic, system-mediated account, which attributes these imbalances to functionally redundant system weights that reduce attention to image and textual inputs. We show that this framework offers a useful empirical perspective on the yes-bias, a common form of hallucination in which VLMs indiscriminately respond `yes'. Causally redistributing attention from the system modality to image and textual inputs substantially suppresses this bias, often outperforming existing approaches. We further present evidence suggesting that system-mediated attention imbalances contribute to the yes-bias by encouraging a default reliance on coarse input representations, which are effective for some tasks but ill-suited to others. Taken together, these findings firmly establish system attention as a key factor in VLM hallucination and highlight its potential as a lever for mitigation.

视觉语言模型幻觉抑制注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。