用医学视觉模型提升放射科报告自动生成的准确性和相关性
Integrating MedCLIP and Cross-Modal Fusion for Automatic Radiology Report Generation
- 用MedCLIP同时提取图像特征和检索报告
- 通过注意力模块融合图文特征,提升报告连贯性
- 在IU-Xray数据集上优于主流方法,适合临床辅助生成
自动化放射科报告生成可显著减轻放射科医生的工作负担,并提高临床文档的准确性、一致性和效率。我们提出一种新的跨模态框架,利用MedCLIP作为视觉提取器和检索机制,改进医学报告生成过程。通过注意力提取模块获取检索到的报告特征与图像特征,并经融合模块整合,提升了生成报告的连贯性和临床相关性。在广泛使用的IU-Xray数据集上的实验结果表明,该方法在报告质量和相关性方面均优于常用方法。消融实验进一步验证了框架有效性,强调了准确报告检索与特征融合在生成全面医学报告中的重要性。
原文摘要 · Abstract (English)
Automating radiology report generation can significantly reduce the workload of radiologists and enhance the accuracy, consistency, and efficiency of clinical documentation.We propose a novel cross-modal framework that uses MedCLIP as both a vision extractor and a retrieval mechanism to improve the process of medical report generation.By extracting retrieved report features and image features through an attention-based extract module, and integrating them with a fusion module, our method improves the coherence and clinical relevance of generated reports.Experimental results on the widely used IU-Xray dataset demonstrate the effectiveness of our approach, showing improvements over commonly used methods in both report quality and relevance.Additionally, ablation studies provide further validation of the framework, highlighting the importance of accurate report retrieval and feature integration in generating comprehensive medical reports.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。