arXiv:2410.16239cs.AIcs.CV2024-10被引 6

融合胸片、心电图与报告,用对比学习提升医疗诊断准确率。

MoRE: Multi-Modal Contrastive Pre-training with Transformers on X-Rays, ECGs, and Diagnostic Report

  • 用Transformer统一编码影像与文本,通过对比损失对齐多模态特征。
  • 在Mimic-IV等4个数据集上达到当前最优性能,零样本分类效果突出。
  • 适合医疗AI研究者,尤其关注多模态融合与可解释性的场景。

本文提出一种新型多模态对比预训练框架MoRE,首次将胸片、心电图(ECG)与放射科/心内科报告进行联合建模。采用Transformer将不同模态编码至统一表示空间,利用对比损失对齐模态特征,支持零样本分类与多模态检索等下游任务。为降低参数量,引入LoRA-Peft微调大语言模型,并在视觉变压器(ViT)中采用线性注意力降噪策略以优化计算效率。我们还提供了新颖的多模态注意力解释与检索能力。在Mimic-IV、CheXpert、Edema Severity和PTBXL四个数据集上,该方法显著超越现有技术,展现出对复杂跨模态关系的捕捉能力与临床鲁棒性,为医疗领域多模态学习提供新范式。

原文摘要 · Abstract (English)

In this paper, we introduce a novel Multi-Modal Contrastive Pre-training Framework that synergistically combines X-rays, electrocardiograms (ECGs), and radiology/cardiology reports. Our approach leverages transformers to encode these diverse modalities into a unified representation space, aiming to enhance diagnostic accuracy and facilitate comprehensive patient assessments. We utilize LoRA-Peft to significantly reduce trainable parameters in the LLM and incorporate recent linear attention dropping strategy in the Vision Transformer(ViT) for smoother attention. Furthermore, we provide novel multimodal attention explanations and retrieval for our model. To the best of our knowledge, we are the first to propose an integrated model that combines X-ray, ECG, and Radiology/Cardiology Report with this approach. By utilizing contrastive loss, MoRE effectively aligns modality-specific features into a coherent embedding, which supports various downstream tasks such as zero-shot classification and multimodal retrieval. Employing our proposed methodology, we achieve state-of-the-art (SOTA) on the Mimic-IV, CheXpert, Edema Severity, and PtbXl downstream datasets, surpassing existing multimodal approaches. Our proposed framework shows significant improvements in capturing intricate inter-modal relationships and its robustness in medical diagnosis that establishes a framework for future research in multimodal learning in the healthcare sector.

多模态医疗AI对比学习影像报告

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。