arXiv:2409.13612cs.CV2024-09中稿 · ACL被引 9

无需标注和大模型,用场景图自动评估视觉语言模型幻觉。

FIHA: Autonomous Hallucination Evaluation in Vision-Language Models with Davidson Scene Graphs

  • 基于戴维森场景图自动生成问答对,无须人工标注。
  • 在MSCOCO和Foggy数据集上构建了首个全方面幻觉评估基准。
  • 可检测不同幻觉类型间的依赖关系,提升评估可靠性。

大型视觉语言模型(LVLMs)快速发展的同时,幻觉问题普遍存在,亟需低成本且全面的评估方法。现有方法多依赖昂贵的人工标注,且难以覆盖关系、属性及跨方面依赖等多维度幻觉。为此,本文提出FIHA(自主细粒度幻觉评估),可在无大模型和无标注条件下评估LVLM幻觉,并建模各类幻觉之间的依赖关系。FIHA可低成本生成任意图像数据集上的问答对,实现从图像与文本双视角进行幻觉评估。基于此,我们构建了名为FIHA-v1的基准数据集,涵盖来自MSCOCO和Foggy数据集的多样化问题。通过戴维森场景图(DSG)组织问答对结构,进一步提升评估可信度。我们使用FIHA-v1评估了多个代表性模型,揭示其在幻觉检测中的局限性与挑战,并公开代码与数据。

原文摘要 · Abstract (English)

The rapid development of Large Vision-Language Models (LVLMs) often comes with widespread hallucination issues, making cost-effective and comprehensive assessments increasingly vital. Current approaches mainly rely on costly annotations and are not comprehensive -- in terms of evaluating all aspects such as relations, attributes, and dependencies between aspects. Therefore, we introduce the FIHA (autonomous Fine-graIned Hallucination evAluation evaluation in LVLMs), which could access hallucination LVLMs in the LLM-free and annotation-free way and model the dependency between different types of hallucinations. FIHA can generate Q&A pairs on any image dataset at minimal cost, enabling hallucination assessment from both image and caption. Based on this approach, we introduce a benchmark called FIHA-v1, which consists of diverse questions on various images from MSCOCO and Foggy. Furthermore, we use the Davidson Scene Graph (DSG) to organize the structure among Q&A pairs, in which we can increase the reliability of the evaluation. We evaluate representative models using FIHA-v1, highlighting their limitations and challenges. We released our code and data.

幻觉评估视觉语言模型场景图自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。