arXiv:2412.12001cs.CLcs.CV2024-12被引 17

让报告生成模型适应多种临床输入,减少幻觉。

LLM-RG4: Flexible and Factual Radiology Report Generation across Diverse Input Contexts

  • 用大模型+自适应融合模块,灵活处理不同输入组合。
  • 在新数据集上实现最佳生成效果,幻觉率显著降低。
  • 适合需要多场景报告生成的医疗AI研发与临床部署。

放射科报告撰写需根据可用信息和临床需求灵活调整内容。然而,现有报告生成(RRG)模型多限定于固定任务范式,如仅从单张图像预测‘发现’部分,导致输入输出不匹配,缺乏灵活性,易产生与输入无关的错误描述。为弥合模型与临床实际需求间的差距,我们构建了MIMIC-RG4数据集,涵盖四种常见报告撰写场景,确保输入输出精准对应。同时提出基于大语言模型的LLM-RG4框架,利用其指令跟随能力与通用知识,设计自适应标记融合模块以支持多样输入,且计算开销低。进一步提出标记级损失加权策略,引导模型关注正向与不确定描述。实验表明,LLM-RG4在MIMIC-RG4与MIMIC-CXR数据集上均达到领先性能,显著降低输入无关幻觉,优于当前开源模型。

原文摘要 · Abstract (English)

Drafting radiology reports is a complex task requiring flexibility, where radiologists tail content to available information and particular clinical demands. However, most current radiology report generation (RRG) models are constrained to a fixed task paradigm, such as predicting the full ``finding'' section from a single image, inherently involving a mismatch between inputs and outputs. The trained models lack the flexibility for diverse inputs and could generate harmful, input-agnostic hallucinations. To bridge the gap between current RRG models and the clinical demands in practice, we first develop a data generation pipeline to create a new MIMIC-RG4 dataset, which considers four common radiology report drafting scenarios and has perfectly corresponded input and output. Secondly, we propose a novel large language model (LLM) based RRG framework, namely LLM-RG4, which utilizes LLM's flexible instruction-following capabilities and extensive general knowledge. We further develop an adaptive token fusion module that offers flexibility to handle diverse scenarios with different input combinations, while minimizing the additional computational burden associated with increased input volumes. Besides, we propose a token-level loss weighting strategy to direct the model's attention towards positive and uncertain descriptions. Experimental results demonstrate that LLM-RG4 achieves state-of-the-art performance in both clinical efficiency and natural language generation on the MIMIC-RG4 and MIMIC-CXR datasets. We quantitatively demonstrate that our model has minimal input-agnostic hallucinations, whereas current open-source models commonly suffer from this problem.

医学报告生成大模型应用幻觉抑制多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。