arXiv:2502.12591cs.CVcs.CL2025-02被引 4

无需大模型推理,用视觉知识库快速检测图文生成中的幻觉

CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base

  • 构建视觉辅助知识库,结合图像与语义关系进行多步验证
  • 在POPE和R-Bench上表现接近顶尖方法,速度提升超10倍
  • 适合需要低成本、离线部署的幻觉检测场景

大型视觉语言模型(LVLM)虽具强大多模态推理能力,但易产生幻觉,尤其在描述中虚构不存在物体或错误属性。现有检测方法性能强,但依赖昂贵API调用与迭代式大模型验证,难以用于大规模或离线场景。为此,我们提出CutPaste&Find——一种轻量级、无需训练的幻觉检测框架。该方法利用现成的视觉与语言模块,通过不依赖LVLM推理的多步验证实现高效检测。核心是包含丰富实体-属性关系及对应图像表征的视觉辅助知识库,并引入缩放因子优化相似度评分,缓解真实图文对间对齐不佳问题。在POPE与R-Bench等基准数据集上的全面评估表明,CutPaste&Find在幻觉检测性能上达到领先水平,同时效率显著高于以往方法。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have demonstrated impressive multimodal reasoning capabilities, but they remain susceptible to hallucination, particularly object hallucination where non-existent objects or incorrect attributes are fabricated in generated descriptions. Existing detection methods achieve strong performance but rely heavily on expensive API calls and iterative LVLM-based validation, making them impractical for large-scale or offline use. To address these limitations, we propose CutPaste\&Find, a lightweight and training-free framework for detecting hallucinations in LVLM-generated outputs. Our approach leverages off-the-shelf visual and linguistic modules to perform multi-step verification efficiently without requiring LVLM inference. At the core of our framework is a Visual-aid Knowledge Base that encodes rich entity-attribute relationships and associated image representations. We introduce a scaling factor to refine similarity scores, mitigating the issue of suboptimal alignment values even for ground-truth image-text pairs. Comprehensive evaluations on benchmark datasets, including POPE and R-Bench, demonstrate that CutPaste\&Find achieves competitive hallucination detection performance while being significantly more efficient and cost-effective than previous methods.

幻觉检测视觉知识库多模态轻量级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。