arXiv:2510.09256cs.CV2025-10被引 2

用语义熵过滤幻觉问题,让医学视觉问答更准。

Hallucination Filtering in Radiology Vision-Language Models Using Discrete Semantic Entropy

  • 通过离散语义熵衡量答案一致性,识别易出错的问题
  • 过滤高熵问题后,准确率从51.7%提升至76.3%
  • 无需模型修改,适合临床部署的黑箱VLM纠错

本研究评估使用离散语义熵(DSE)过滤可能引发幻觉的问题,能否提升黑箱视觉语言模型(VLM)在放射影像视觉问答(VQA)中的准确性。回顾性分析两个公开脱敏数据集:VQA-Med 2019基准(500张图像,含临床问题与短文本答案)和诊断放射学数据集(206例:60次CT、60次MRI、60次X光、26次血管造影),对应真实诊断。GPT-4o与GPT-4.1分别以温度1.0运行15次,基线准确率基于温度0.1结果计算。通过双向蕴含检查合并语义等价回答,基于语义聚类频率计算DSE。在排除DSE > 0.6或 > 0.3的问题后重新计算准确率。采用自助抽样法获取p值与95%置信区间,多重比较校正阈值设为p < .004。共706个图像-问题对中,GPT-4o基线准确率为51.7%,GPT-4.1为54.8%。过滤高熵问题(DSE > 0.3)后,剩余问题准确率分别为76.3%(保留334/706)和63.8%(保留499/706),均显著提升(p < .001)。两种数据集上均有增益,且经校正后仍显著。该方法通过量化语义不一致实现可靠幻觉检测,显著提升诊断答案准确率,为临床VLM应用提供有效过滤策略。

原文摘要 · Abstract (English)

To determine whether using discrete semantic entropy (DSE) to reject questions likely to generate hallucinations can improve the accuracy of black-box vision-language models (VLMs) in radiologic image based visual question answering (VQA). This retrospective study evaluated DSE using two publicly available, de-identified datasets: the VQA-Med 2019 benchmark (500 images with clinical questions and short-text answers) and a diagnostic radiology dataset (206 cases: 60 computed tomography scans, 60 magnetic resonance images, 60 radiographs, 26 angiograms) with corresponding ground-truth diagnoses. GPT-4o and GPT-4.1 (Generative Pretrained Transformer; OpenAI) answered each question 15 times using a temperature of 1.0. Baseline accuracy was determined using low-temperature answers (temperature 0.1). Meaning-equivalent responses were grouped using bidirectional entailment checks, and DSE was computed from the relative frequencies of the resulting semantic clusters. Accuracy was recalculated after excluding questions with DSE > 0.6 or > 0.3. p-values and 95% confidence intervals were obtained using bootstrap resampling and a Bonferroni-corrected threshold of p < .004 for statistical significance. Across 706 image-question pairs, baseline accuracy was 51.7% for GPT-4o and 54.8% for GPT-4.1. After filtering out high-entropy questions (DSE > 0.3), accuracy on the remaining questions was 76.3% (retained questions: 334/706) for GPT-4o and 63.8% (retained questions: 499/706) for GPT-4.1 (both p < .001). Accuracy gains were observed across both datasets and largely remained statistically significant after Bonferroni correction. DSE enables reliable hallucination detection in black-box VLMs by quantifying semantic inconsistency. This method significantly improves diagnostic answer accuracy and offers a filtering strategy for clinical VLM applications.

医学视觉问答幻觉检测语义熵VLM纠错

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。