arXiv:2503.13107cs.CVcs.AI2025-03CVPR被引 45

通过增强视觉信号提升模型对图像的注意力,有效减少幻觉且不降速

ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models

  • 在中间层融合时强化视觉特征,减少语言主导倾向
  • 多模型测试显示幻觉显著降低,生成内容保持准确连贯
  • 无需额外训练,即插即用,推理速度不变

对比解码策略广泛用于缓解多模态大语言模型(MLLMs)中的物体幻觉问题。通过降低对语言先验的依赖,这类方法确保生成内容紧贴视觉输入,实现上下文准确输出。由于对比解码无需额外训练或外部工具,兼具计算高效与通用性,极具吸引力。然而,现有方法存在两大局限:(1) 一味压制语言先验会损害生成内容的连贯性与准确性;(2) 处理对比输入增加计算负担,显著降低推理速度。为此,我们提出视觉增强融合(VAF),一种即插即用技术,在模型中层强化视觉信号关注,该处为模态融合主要发生位置。此方法可更有效地捕捉视觉特征,减轻模型对语言模态的偏倚。实验结果表明,VAF在不影响推理速度的前提下,显著降低多种MLLMs的幻觉,同时保持生成内容的连贯性与准确性。

原文摘要 · Abstract (English)

Contrastive decoding strategies are widely used to mitigate object hallucinations in multimodal large language models (MLLMs). By reducing over-reliance on language priors, these strategies ensure that generated content remains closely grounded in visual inputs, producing contextually accurate outputs. Since contrastive decoding requires no additional training or external tools, it offers both computational efficiency and versatility, making it highly attractive. However, these methods present two main limitations: (1) bluntly suppressing language priors can compromise coherence and accuracy of generated content, and (2) processing contrastive inputs adds computational load, significantly slowing inference speed. To address these challenges, we propose Visual Amplification Fusion (VAF), a plug-and-play technique that enhances attention to visual signals within the model's middle layers, where modality fusion predominantly occurs. This approach enables more effective capture of visual features, reducing the model's bias toward language modality. Experimental results demonstrate that VAF significantly reduces hallucinations across various MLLMs without affecting inference speed, while maintaining coherence and accuracy in generated outputs.

多模态幻觉抑制视觉增强即插即用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。