arXiv:2605.10676cs.CVcs.LG2026-05被引 1

通过对抗性扰动恢复视觉语言平衡,提升大模型推理可信度

Not Blind but Silenced: Rebalancing Vision and Language via Adversarial Counter-Commonsense Equilibrium

论文配图:Not Blind but Silenced: Rebalancing Vision and Language via Adversarial Counter-Commonsense Equilibrium
图 1 · 摘自论文原文
  • 用反常识补丁扰动视觉上下文,识别并抑制幻觉源头
  • 在不增加计算开销前提下,显著降低无关注意力集中现象
  • 适用于所有无需训练的视觉语言模型,特别适合高可靠性场景

在多模态大模型解码过程中,注意力常异常聚焦于无关图像标记。现有方法将此类现象视为无效噪声并强制引导注意力聚焦关键视觉信息,但我们认为这些标记承载着重要的视觉与叙事逻辑,强制纠正反而加剧了视觉-语言失衡。从‘解码即博弈’视角出发,我们揭示幻觉源于语言先验与视觉信息间的均衡失调。提出无需训练的对抗性反常识均衡(ACE)框架,通过反常识补丁扰动视觉上下文,利用真实视觉特征在扰动下保持稳定而幻觉信号波动剧烈的特性,实现动态博弈式解码。该策略精准抑制对扰动敏感的语言先验,同时补偿稳定的视觉信号,重建平衡。大量实验表明,作为即插即用方案,ACE在几乎无额外推理开销下显著提升模型可信度。

原文摘要 · Abstract (English)

During MLLM decoding, attention often abnormally concentrates on irrelevant image tokens. While existing research dismisses this as invalid noise and forcibly redirects attention to compel focusing on key image information, we argue these tokens are critical carriers of visual and narrative logic, and such coercive corrections exacerbate visual-language imbalance. Adopting a "decoding-as-game" perspective, we reveal that hallucinations stem from an equilibrium imbalance between linguistic priors and visual information. We propose Adversarial Counter-Commonsense Equilibrium (ACE), a training-free framework that perturbs visual context via counter-commonsense patches. Leveraging the fact that authentic visual features remain stable under perturbation while hallucinations fluctuate, ACE implements a dynamic game decoding strategy. This approach precisely suppresses perturbation-sensitive priors while compensating for stable visual signals to restore balance. Extensive experiments demonstrate that ACE, as a plug-and-play strategy, enhances model trustworthiness with negligible inference overhead.

视觉语言幻觉抑制解码机制对抗训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。