首个胃肠道内镜视觉问答数据集上的幻觉检测基准,揭示白盒方法显著优于黑盒。
A Benchmark for Hallucination Detection in VLMs for Gastrointestinal Endoscopy

- 构建胃肠道内镜VQA数据集Gut-VLM,含4392个测试样本
- 白盒方法ReXTrust在5个模型上均达最高AUC(最高93.0)
- 发现高置信度编造是主流方法的系统性失败点
视觉语言模型(VLMs)易产生幻觉,严重阻碍其在临床中的安全应用。现有幻觉检测方法多基于放射科数据集如MIMIC-CXR和VQA-RAD评估,而胃肠道(GI)内镜领域仍缺乏研究。本文在包含4,392个测试VQA对的Gut-VLM数据集上,对九种幻觉检测方法进行基准测试,覆盖五个VLM模型:MedGemma-4B、MedGemma-27B、LLaVA-Med-7B、LLaVA-v1.6-7B和Lingshu-32B。方法涵盖三类:黑盒(RadFlag、SelfCheckGPT-NLI)、灰盒(AvgProb、AvgEnt、MaxProb、MaxEnt、Semantic Entropy、VASE)及白盒(ReXTrust)。结果表明,白盒方法ReXTrust在所有模型中均取得最高AUC,显著优于各模型最强替代方法(配对置换检验,p < 0.001),在MedGemma-4B上达到峰值AUC 93.0。白盒隐藏状态访问带来平均19.5 AUC点优势(范围9.5–33.5),即使在LLaVA-v1.6-7B上也保持良好表现(AUC 79.9),而黑盒与聚类型灰盒方法性能接近随机。非白盒方法中,基于词元级别的灰盒统计(MaxEnt、MaxProb)表现最佳,优于聚类型灰盒(Semantic Entropy、VASE)和黑盒方法。进一步发现‘高置信度编造’——模型以高一致性或高词元概率生成错误内容——是统一性与不确定性方法的系统性失效模式。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are prone to hallucination, which remains a major barrier to their safe deployment in clinical practice. To date, most hallucination detection methods have been evaluated on radiology benchmarks such as MIMIC-CXR and VQA-RAD, while gastrointestinal (GI) endoscopy remains largely underexplored. In this paper, we benchmark nine hallucination detection methods on the Gut-VLM dataset, a GI diagnostic Visual Question Answering (VQA) dataset with 4,392 test VQA pairs, across five VLMs (MedGemma-4B, MedGemma-27B, LLaVA-Med-7B, LLaVA-v1.6-7B, and Lingshu-32B). The methods span three categories: black-box methods (RadFlag, SelfCheckGPT-NLI), gray-box methods (AvgProb, AvgEnt, MaxProb, MaxEnt, Semantic Entropy, and VASE), and a white-box method (ReXTrust). Our results show that ReXTrust, a white-box method, achieves the highest AUC across all five models, outperforming the strongest alternative method on each VLM by a statistically significant margin (paired permutation test, p < 0.001 in all cases), reaching a peak AUC of 93.0 on MedGemma-4B. White-box hidden-state access provides a consistent advantage of 19.5 AUC points on average (range: 9.5--33.5), with ReXTrust maintaining strong performance even on LLaVA-v1.6-7B (AUC 79.9), where black-box methods and clustering-based gray-box methods collapse to near-chance performance. Among non-white-box methods, token-level gray-box statistics (MaxEnt, MaxProb) are the strongest alternatives, outperforming both clustering-based gray-box methods (Semantic Entropy, VASE) and black-box approaches on average. We further identify confident confabulation, a failure mode in which models hallucinate with high inter-sample consistency or high token-level probability, as a systemic failure for both consistency and uncertainty-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。