用跨模态融合与医学预训练提升胸部X光报告生成准确率
EIR: Enhanced Image Representations for Medical Report Generation
- 通过跨模态变换器融合医学元数据与图像特征,解决异构信息不对称问题
- 在MIMIC和Open-I数据集上报告生成指标显著优于现有方法
- 适合需要高精度医学影像报告自动化的临床研究与系统开发
从胸部X光图像生成医学报告是放射科医生在紧急情况下的一项关键且耗时的任务。为减轻放射科医生压力并降低误诊风险,近年来大量研究致力于自动医学报告生成。现有方法通常利用患者临床病历或相似病例报告构建的医学图谱等元数据来增强图像表示,但这些方法仅通过简单的“相加+层归一化”操作融合元数据与视觉特征,导致二者分布差异引发的信息不对称问题。此外,胸部X光图像多采用基于自然图像预训练的模型编码,存在显著的通用领域与医学领域之间的域差距。为此,我们提出一种新方法——增强图像表示(EIR),通过跨模态变换器融合元数据与图像表示,有效缓解信息不对称;同时采用医学领域预训练模型编码医学图像,有效弥合图像表征的域差距。在广泛使用的MIMIC和Open-I数据集上的实验结果验证了该方法的有效性。
原文摘要 · Abstract (English)
Generating medical reports from chest X-ray images is a critical and time-consuming task for radiologists, especially in emergencies. To alleviate the stress on radiologists and reduce the risk of misdiagnosis, numerous research efforts have been dedicated to automatic medical report generation in recent years. Most recent studies have developed methods that represent images by utilizing various medical metadata, such as the clinical document history of the current patient and the medical graphs constructed from retrieved reports of other similar patients. However, all existing methods integrate additional metadata representations with visual representations through a simple "Add and LayerNorm" operation, which suffers from the information asymmetry problem due to the distinct distributions between them. In addition, chest X-ray images are usually represented using pre-trained models based on natural domain images, which exhibit an obvious domain gap between general and medical domain images. To this end, we propose a novel approach called Enhanced Image Representations (EIR) for generating accurate chest X-ray reports. We utilize cross-modal transformers to fuse metadata representations with image representations, thereby effectively addressing the information asymmetry problem between them, and we leverage medical domain pre-trained models to encode medical images, effectively bridging the domain gap for image representation. Experimental results on the widely used MIMIC and Open-I datasets demonstrate the effectiveness of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。