让医学影像报告更准:结合模型内知识与外部信息
RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge Injection
- 先提取模型已有知识,再补充外部信息
- 在三个数据集上报告准确率超越现有模型
- 适合医疗AI研究者和临床辅助系统开发者
大型语言模型(LLMs)在多个领域展现出强大能力,包括医学影像报告生成。以往方法尝试使用多模态LLM,并通过引入领域特定知识检索来提升性能,但常忽略模型内部已有的知识,造成信息冗余。为此,我们提出RADAR框架,通过系统性融合模型内知识与外部检索信息,提升报告生成质量。具体而言,首先提取与专家图像分类结果一致的模型内部知识;随后检索相关补充知识以进一步丰富信息;最后整合两源信息生成更准确、更丰富的放射科报告。在MIMIC-CXR、CheXpert-Plus和IU X-ray数据集上的大量实验表明,本模型在语言质量和临床准确性方面均优于当前最先进的LLM。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated remarkable capabilities in various domains, including radiology report generation. Previous approaches have attempted to utilize multimodal LLMs for this task, enhancing their performance through the integration of domain-specific knowledge retrieval. However, these approaches often overlook the knowledge already embedded within the LLMs, leading to redundant information integration. To address this limitation, we propose Radar, a framework for enhancing radiology report generation with supplementary knowledge injection. Radar improves report generation by systematically leveraging both the internal knowledge of an LLM and externally retrieved information. Specifically, it first extracts the model's acquired knowledge that aligns with expert image-based classification outputs. It then retrieves relevant supplementary knowledge to further enrich this information. Finally, by aggregating both sources, Radar generates more accurate and informative radiology reports. Extensive experiments on MIMIC-CXR, CheXpert-Plus, and IU X-ray demonstrate that our model outperforms state-of-the-art LLMs in both language quality and clinical accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。