arXiv:2412.15264cs.CLcs.AI2024-12中稿 · AIMedHealth 10 pag…被引 15

用视觉语言模型隐藏状态检测影像报告幻觉,提升医疗AI安全

ReXTrust: A Model for Fine-Grained Hallucination Detection in AI-Generated Radiology Reports

  • 通过大模型隐藏状态生成逐发现的幻觉风险评分
  • 在MIMIC-CXR上实现0.8751的全发现AUROC,关键发现达0.8963
  • 适合关注医疗AI安全与可解释性的研究者和临床开发者

AI生成影像报告的广泛应用亟需可靠的幻觉检测方法——即识别可能影响患者诊疗的虚假或无依据陈述。本文提出ReXTrust框架,利用大规模视觉-语言模型的隐藏状态序列,对生成报告中的每个医学发现进行细粒度幻觉风险评分。我们在MIMIC-CXR数据集的一个子集上评估该方法,结果表明其性能优于现有技术:所有发现的AUROC达到0.8751,临床重要发现的AUROC更达0.8963。实验表明,基于模型内部状态的白盒方法可为医疗AI系统提供可靠的幻觉检测能力,有助于提升自动化影像报告的安全性与可信度。

原文摘要 · Abstract (English)

The increasing adoption of AI-generated radiology reports necessitates robust methods for detecting hallucinations--false or unfounded statements that could impact patient care. We present ReXTrust, a novel framework for fine-grained hallucination detection in AI-generated radiology reports. Our approach leverages sequences of hidden states from large vision-language models to produce finding-level hallucination risk scores. We evaluate ReXTrust on a subset of the MIMIC-CXR dataset and demonstrate superior performance compared to existing approaches, achieving an AUROC of 0.8751 across all findings and 0.8963 on clinically significant findings. Our results show that white-box approaches leveraging model hidden states can provide reliable hallucination detection for medical AI systems, potentially improving the safety and reliability of automated radiology reporting.

医疗AI幻觉检测视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。