arXiv:2602.11737cs.CVcs.CL2026-02Conference of the …被引 3

通过视觉对比解码削弱物体幻觉,提升多模态模型准确性

Mask What Matters: Mitigating Object Hallucinations in Multimodal Large Language Models with Object-Aligned Visual Contrastive Decoding

  • 用物体对齐辅助视图,移除显著视觉特征增强对比信号
  • 在两个基准上显著降低物体幻觉,适配多种大模型
  • 无需修改原模型,单次前向计算即可接入,效率高

我们研究多模态大语言模型中的物体幻觉问题,改进视觉对比解码(VCD)方法,通过构建物体对齐的辅助视图实现。利用自监督视觉变换器中的物体中心注意力机制,主动移除最显著的视觉证据,以干扰不支持的生成词元,并产生更强的对比信号。该方法与提示无关、与模型无关,可无缝嵌入现有VCD流程,仅需一次可缓存的前向传播,计算开销极小。实验证明,该方法在两个主流物体幻觉基准上,对两种MLLM均表现出一致的性能提升。

原文摘要 · Abstract (English)

We study object hallucination in Multimodal Large Language Models (MLLMs) and improve visual contrastive decoding (VCD) by constructing an object-aligned auxiliary view. We leverage object-centric attention in self-supervised Vision Transformers. In particular, we remove the most salient visual evidence to construct an auxiliary view that disrupts unsupported tokens and produces a stronger contrast signal. Our method is prompt-agnostic, model-agnostic, and can be seamlessly plugged into the existing VCD pipeline with little computation overhead, i.e., a single cacheable forward pass. Empirically, our method demonstrates consistent gains on two popular object hallucination benchmarks across two MLLMs.

多模态幻觉抑制视觉对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。