用概念反事实解释图像描述中的幻觉,让错误原因可懂可改。
HalCECE: A Framework for Explainable Hallucination Detection through Conceptual Counterfactuals in Image Captioning
- 基于概念反事实生成语义最小修改建议,黑盒推导非幻觉描述。
- 通过分层分解幻觉概念,实现对幻觉成因的深入分析。
- 首次关注角色幻觉,适合研究模型可信性与可解释性的学者。
在人工智能快速发展背景下,视觉语言模型中的幻觉现象成为关键研究方向。本文针对广泛使用的图像描述模型中存在的幻觉问题,引入概念反事实解释技术进行分析。所采用的确定性且高效的反事实框架,能够基于分层知识生成语义最小的编辑建议,以黑箱方式将幻觉描述转换为非幻觉版本。提出的HalCECE框架具有高度可解释性,不仅提供有意义的文本修改建议,还通过幻觉概念的分层分解实现全面分析。另一创新点在于首次探讨角色幻觉,揭示视觉概念间关联在幻觉检测中的作用。整体上,该工作为视觉语言模型幻觉检测提供了可解释的新路径,有助于当前及未来系统的可信评估。
原文摘要 · Abstract (English)
In the dynamic landscape of artificial intelligence, the exploration of hallucinations within vision-language (VL) models emerges as a critical frontier. This work delves into the intricacies of hallucinatory phenomena exhibited by widely used image captioners, unraveling interesting patterns. Specifically, we step upon previously introduced techniques of conceptual counterfactual explanations to address VL hallucinations. The deterministic and efficient nature of the employed conceptual counterfactuals backbone is able to suggest semantically minimal edits driven by hierarchical knowledge, so that the transition from a hallucinated caption to a non-hallucinated one is performed in a black-box manner. HalCECE, our proposed hallucination detection framework is highly interpretable, by providing semantically meaningful edits apart from standalone numbers, while the hierarchical decomposition of hallucinated concepts leads to a thorough hallucination analysis. Another novelty tied to the current work is the investigation of role hallucinations, being one of the first works to involve interconnections between visual concepts in hallucination detection. Overall, HalCECE recommends an explainable direction to the crucial field of VL hallucination detection, thus fostering trustworthy evaluation of current and future VL systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。