arXiv:2507.19024cs.CV2025-07IJCV综述被引 30

系统梳理多模态幻觉的评估与检测方法,助力模型更真实可信。

A Survey of Multimodal Hallucination Evaluation and Detection

  • 按忠实度与事实性对幻觉分类,构建清晰分析框架。
  • 汇总图文生成任务中主流评估基准与检测技术。
  • 适合关注AI生成内容真实性的研究者与开发者。

多模态大语言模型(MLLMs)已成为融合视觉与文本信息的强大范式,支持多种多模态任务。然而,这些模型常出现幻觉,生成看似合理但与输入内容或已知世界知识矛盾的内容。本综述深入分析了图像到文本(I2T)和文本到图像(T2I)生成任务中的幻觉评估基准与检测方法。首先,基于忠实度与事实性提出幻觉分类体系,涵盖实践中常见的幻觉类型。随后,概述现有T2I与I2T任务的评估基准,重点说明其构建过程、评估目标与采用的指标。进一步总结近年幻觉检测方法进展,旨在实例级识别幻觉内容,作为基准评估的实用补充。最后,指出当前基准与检测方法的关键局限,并提出未来研究方向。

原文摘要 · Abstract (English)

Multi-modal Large Language Models (MLLMs) have emerged as a powerful paradigm for integrating visual and textual information, supporting a wide range of multi-modal tasks. However, these models often suffer from hallucination, producing content that appears plausible but contradicts the input content or established world knowledge. This survey offers an in-depth review of hallucination evaluation benchmarks and detection methods across Image-to-Text (I2T) and Text-to-image (T2I) generation tasks. Specifically, we first propose a taxonomy of hallucination based on faithfulness and factuality, incorporating the common types of hallucinations observed in practice. Then we provide an overview of existing hallucination evaluation benchmarks for both T2I and I2T tasks, highlighting their construction process, evaluation objectives, and employed metrics. Furthermore, we summarize recent advances in hallucination detection methods, which aims to identify hallucinated content at the instance level and serve as a practical complement of benchmark-based evaluation. Finally, we highlight key limitations in current benchmarks and detection methods, and outline potential directions for future research.

多模态幻觉检测评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。