arXiv:2601.15745cs.CL2026-01被引 2

用知识增强与细粒度奖励减少医学报告生成中的幻觉

Hallucination Mitigating for Medical Report Generation

  • 引入医学知识库和净化模块,提升输入信息的准确性
  • 在IU-Xray和MIMIC-CXR上显著降低幻觉率,提升报告质量
  • 适合医疗AI研发、临床辅助系统开发人员参考

在医学报告生成(MRG)领域,自然语言处理技术已成为减轻放射科医生工作负担的重要工具。尽管大视觉语言模型(LVLMs)在理解自然语言方面表现出色,但其容易生成看似合理实则错误的描述(即“幻觉”),这在医疗这一高度敏感且关键的领域引发担忧。本文提出一种新框架KERM(知识增强的细粒度强化奖励医学报告生成),通过MedCLIP从精心构建的知识语料库中检索相关病灶事实句,再经净化模块确保知识与患者临床背景一致,最后采用细粒度奖励引导模型生成更具支持性与临床相关性的描述,实现输出与理想行为对齐。在IU-Xray和MIMIC-CXR数据集上的实验验证了该方法在抑制幻觉和提升报告质量方面的有效性。

原文摘要 · Abstract (English)

In the realm of medical report generation (MRG), the integration of natural language processing has emerged as a vital tool to alleviate the workload of radiologists. Despite the impressive capabilities demonstrated by large vision language models (LVLMs) in understanding natural language, their susceptibility to generating plausible yet inaccurate claims, known as ``hallucinations'', raises concerns-especially in the nuanced and critical field of medical. In this work, we introduce a framework, \textbf{K}nowledge-\textbf{E}nhanced with Fine-Grained \textbf{R}einforced Rewards \textbf{M}edical Report Generation (KERM), to tackle the issue. Our approach refines the input to the LVLM by first utilizing MedCLIP for knowledge retrieval, incorporating relevant lesion fact sentences from a curated knowledge corpus. We then introduce a novel purification module to ensure the retrieved knowledge is contextually relevant to the patient's clinical context. Subsequently, we employ fine-grained rewards to guide these models in generating highly supportive and clinically relevant descriptions, ensuring the alignment of model's outputs with desired behaviors. Experimental results on IU-Xray and MIMIC-CXR datasets validate the effectiveness of our approach in mitigating hallucinations and enhancing report quality.

医学报告生成幻觉抑制视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。