通过全局与局部相似性联合检测大模型的物体幻觉问题
GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
- 融合图像与文本的全局和局部嵌入相似性进行检测
- 在多个场景下优于现有基线方法,检测性能显著提升
- 无需训练,适用于多种实际应用中的幻觉识别
大型视觉语言模型中的物体幻觉问题严重威胁其在真实场景中的安全部署。现有方法通常仅从全局或局部视角出发,可能影响检测可靠性。本文提出GLSim,一种无需训练的物体幻觉检测框架,通过融合图像与文本模态间的互补全局与局部嵌入相似性信号,在多样化场景中实现更准确、可靠的幻觉检测。我们全面评估了现有方法,结果表明GLSim在多个基准上显著超越竞争基线,展现出优越的检测性能。
原文摘要 · Abstract (English)
Object hallucination in large vision-language models presents a significant challenge to their safe deployment in real-world applications. Recent works have proposed object-level hallucination scores to estimate the likelihood of object hallucination; however, these methods typically adopt either a global or local perspective in isolation, which may limit detection reliability. In this paper, we introduce GLSim, a novel training-free object hallucination detection framework that leverages complementary global and local embedding similarity signals between image and text modalities, enabling more accurate and reliable hallucination detection in diverse scenarios. We comprehensively benchmark existing object hallucination detection methods and demonstrate that GLSim achieves superior detection performance, outperforming competitive baselines by a significant margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。