针对3D MRI多器官异常检测,提出模态感知预训练框架,提升视觉语言对齐效果。
3D Modality-Aware Pre-training for Vision-Language Model in MRI Multi-organ Abnormality Detection

- 设计模态感知编码器,隐式学习多模态分布,增强图像与文本对齐
- 在7,392对3D MRI影像报告上训练,显著优于现有视觉语言模型
- 适用于多模态医学影像分析,特别适合放射科医生辅助诊断
视觉语言模型(VLMs)在医学影像复杂诊断任务中展现出巨大潜力。然而,将其应用于多器官医学影像时面临两大挑战:(1)模态特异性视觉-语言对齐,(2)跨模态特征融合。本文提出MedMAP——一种用于3D MRI的医学模态感知预训练框架,通过模态感知视觉-语言对齐阶段和下游多器官异常检测微调阶段,增强视觉-语言表征学习能力。预训练阶段,模态感知编码器隐式捕捉联合模态分布,提升视觉与文本表示间的对齐。随后,冻结文本编码器,仅微调视觉编码器以适应下游任务。为此,我们构建了包含7,392组3D MRI影像-报告对的MedMoM-MRI3D数据集,涵盖十二种MRI模态和九类异常,专为多种3D医学分析任务设计。在该数据集上的大量实验表明,MedMAP在3D MRI多器官异常检测任务中显著优于现有VLMs。代码已开源。
原文摘要 · Abstract (English)
Vision-language models (VLMs) show strong potential for complex diagnostic tasks in medical imaging. However, applying VLMs to multi-organ medical imaging introduces two principal challenges: (1) modality-specific vision-language alignment and (2) cross-modal feature fusion. In this work, we propose MedMAP, a Medical Modality-Aware Pretraining framework that enhances vision-language representation learning in 3D MRI. MedMAP comprises a modality-aware vision-language alignment stage and a fine-tuning stage for multi-organ abnormality detection. During the pre-training stage, the modality-aware encoders implicitly capture the joint modality distribution and improve alignment between visual and textual representations. We then fine-tune the pre-trained vision encoders (while keeping the text encoder frozen) for downstream tasks. To this end, we curated MedMoM-MRI3D, comprising 7,392 3D MRI volume-report pairs spanning twelve MRI modalities and nine abnormalities tailored for various 3D medical analysis tasks. Extensive experiments on MedMoM-MRI3D demonstrate that MedMAP significantly outperforms existing VLMs in 3D MRI-based multi-organ abnormality detection. Our code is available at https://github.com/RomantiDr/MedMAP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。