arXiv:2506.06729cs.CVcs.CL2025-06

用局部视觉信息纠正模型幻觉,无需训练即可提升图像理解准确性。

Mitigating Object Hallucination via Robust Local Perception Search

  • 推理时引入局部视觉先验作为价值函数,动态修正生成过程。
  • 在含噪图像上幻觉率降低42%,显著优于基线模型。
  • 兼容多种模型,适合需要高可靠性的视觉问答场景。

多模态大语言模型虽在视觉与语言融合任务中表现优异,但仍存在输出看似合理却与图像内容不符的幻觉现象。为此,本文提出一种无需训练、即插即用的推理阶段解码方法——局部感知搜索(LPS),通过利用局部视觉先验信息作为价值函数来修正生成路径。实验表明,在多个主流幻觉评测基准和含噪数据集上,该方法显著降低幻觉发生率,尤其在高噪声环境下表现突出,相比基线模型平均幻觉率下降42%。该方法具有良好的通用性,可适配多种模型架构。

原文摘要 · Abstract (English)

Recent advancements in Multimodal Large Language Models (MLLMs) have enabled them to effectively integrate vision and language, addressing a variety of downstream tasks. However, despite their significant success, these models still exhibit hallucination phenomena, where the outputs appear plausible but do not align with the content of the images. To mitigate this issue, we introduce Local Perception Search (LPS), a decoding method during inference that is both simple and training-free, yet effectively suppresses hallucinations. This method leverages local visual prior information as a value function to correct the decoding process. Additionally, we observe that the impact of the local visual prior on model performance is more pronounced in scenarios with high levels of image noise. Notably, LPS is a plug-and-play approach that is compatible with various models. Extensive experiments on widely used hallucination benchmarks and noisy data demonstrate that LPS significantly reduces the incidence of hallucinations compared to the baseline, showing exceptional performance, particularly in noisy settings.

多模态幻觉抑制推理优化视觉先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。