用临床对比解码减少医学影像大模型的幻觉,提升报告准确性
CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding
- 引入双阶段对比机制,在生成时优化词元概率
- 在MIMIC-CXR上使RadGraph-F1最高提升17%
- 无需训练或检索,适合医疗AI落地场景
多模态大语言模型在放射学中通过结合视觉感知与自然语言理解取得显著进展,但常生成缺乏临床依据的描述,即医学幻觉,严重威胁医疗应用的准确性。实证分析发现,提示词引发的幻觉仍普遍存在于放射学多模态模型中,主要源于对临床部分过度敏感。为此,我们提出临床对比解码(CCD),一种无需训练、无需检索的推理框架,利用特定任务的放射科专家模型提供的结构化临床信号。CCD通过双阶段对比机制在生成过程中细化词元级逻辑,从而在不修改基础多模态大模型的前提下提升临床一致性。在三个数据集和多个模型上的实验表明,CCD在放射学报告生成任务中持续提升性能。在MIMIC-CXR数据集上,应用于先进报告生成模型时,RadGraph-F1最高提升17%。该方法为减轻医学幻觉提供了轻量且可泛化的解决方案,有效连接专家模型与多模态大模型在放射学中的应用。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have recently achieved remarkable progress in radiology by integrating visual perception with natural language understanding. However, they often generate clinically unsupported descriptions, known as medical hallucinations, which pose serious risks in medical applications that demand accuracy and image-grounded outputs. Through empirical analysis, we find that prompt-induced hallucinations remain prevalent in radiology MLLMs, largely due to over-sensitivity to clinical sections. To address this, we introduce Clinical Contrastive Decoding (CCD), a training-free and retrieval-free inference framework that integrates structured clinical signals from task-specific radiology expert models. CCD introduces a dual-stage contrastive mechanism to refine token-level logits during generation, thereby enhancing clinical fidelity without modifying the base MLLM. Experiments on three datasets and multiple models demonstrate that CCD consistently improves overall performance on radiology report generation (RRG). On the MIMIC-CXR dataset, it yields up to a 17% improvement in RadGraph-F1 when applied to state-of-the-art RRG models. Our approach provides a lightweight and generalisable solution for mitigating medical hallucinations, effectively bridging expert models and MLLMs in radiology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。