arXiv:2502.20780cs.AIcs.CL2025-02被引 8

构建医疗视觉语言模型幻觉评估基准,提升诊断可靠性

MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models

  • 设计包含百万级指令对的医疗幻觉评测集
  • 微调后模型幻觉率下降,零样本问答性能提升
  • 适合医疗AI研发与临床部署者参考

视觉语言模型在医疗应用中日益普及,但其生成看似合理实则错误的内容(即幻觉)可能危及临床决策。本文提出MedHallTune,一个大规模基准,用于评估和缓解医疗VLM中的幻觉问题。该基准包含超过10万张图像和100万条指令对,涵盖有幻觉与无幻觉样本,并配有真实标注。我们对当前主流医疗及通用VLM进行了全面评估,考察临床准确性、相关性、细节程度和风险等级等关键指标。实验结果表明,使用MedHallTune进行微调可有效降低多个现有模型的幻觉率,显著提升其在下游视觉问答任务上的零样本表现,增强其在实际医疗场景中的可靠性。代码与数据集将开源。

原文摘要 · Abstract (English)

The increasing use of vision-language models (VLMs) in healthcare applications presents great challenges related to hallucinations, in which the models may generate seemingly plausible results that are in fact incorrect. Such hallucinations can jeopardize clinical decision making, potentially harming the diagnosis and treatments. In this work, we propose MedHallTune, a large-scale benchmark designed specifically to evaluate and mitigate hallucinations in medical VLMs. Comprising over 100,000 images and 1,000,000 instruction pairs, MedHallTune includes both hallucination and non-hallucination samples, each with ground-truth annotations. We conduct a comprehensive evaluation of current medical and general VLMs using MedHallTune, assessing their performance across key metrics, including clinical accuracy, relevance, detail level, and risk level. The experimental results show that fine-tuning with MedHallTune successfully improves the ability of several existing models to manage hallucinations and boost their zero-shot performance on downstream visual-question-answering (VQA) tasks, making them more reliable for practical medical applications. Our work contributes to the development of more trustworthy VLMs. Codes and dataset will be available at \href{https://github.com/russellyq/MedHallTune}{MedHallTune}.

医疗AI幻觉检测视觉语言模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。