通过多频扰动减少视觉语言模型幻觉
Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
- 利用图像的高低频特征扰动视觉表示,抑制冗余频率信息
- 在多个模型架构上显著降低物体幻觉率,提升生成真实性
- 方法轻量可插拔,适合与推理阶段方法结合使用
近期,多模态大语言模型(MLLMs)在视觉-语言任务中表现出色。然而,其生成结果常因物体幻觉而失真。我们发现,幻觉的主要原因是模型对图像特定频率特征过度敏感。本文提出多频扰动(MFP)方法,通过同时利用图像的低频和高频特征,在推理时扰动视觉特征表示,并显式抑制冗余的频域特征,从而缓解幻觉问题。实验表明,该方法在多种模型架构上均能有效减少物体幻觉。此外,作为训练阶段方法,MFP可与推理阶段方法结合,在CHAIR基准上达到当前最优性能。
原文摘要 · Abstract (English)
Recently, multimodal large language models (MLLMs) have demonstrated remarkable performance in visual-language tasks. However, the authenticity of the responses generated by MLLMs is often compromised by object hallucinations. We identify that a key cause of these hallucinations is the model's over-susceptibility to specific image frequency features in detecting objects. In this paper, we introduce Multi-Frequency Perturbations (MFP), a simple, cost-effective, and pluggable method that leverages both low-frequency and high-frequency features of images to perturb visual feature representations and explicitly suppress redundant frequency-domain features during inference, thereby mitigating hallucinations. Experimental results demonstrate that our method significantly mitigates object hallucinations across various model architectures. Furthermore, as a training-time method, MFP can be combined with inference-time methods to achieve state-of-the-art performance on the CHAIR benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。