arXiv:2409.02418cs.CV2024-09被引 7

用医学报告指导多器官分割,提升小样本场景下的精度。

MOSMOS: Multi-organ segmentation facilitated by medical report supervision

  • 利用图像与报告的对比学习对齐语义,实现跨模态预训练。
  • 通过多标签识别隐式建立像素与器官标签的对应关系。
  • 可适配2D/3D分割网络,适用于多种疾病和影像模态。

现代医疗系统中存在大量多模态数据(如医学影像与报告),医学视觉语言预训练(Med-VLP)在粗粒度下游任务(如分类、检索、视觉问答)中表现卓越。然而,将Med-VLP学到的知识迁移至细粒度的多器官分割任务仍鲜有研究。多器官分割挑战主要源于缺乏大规模全标注数据集,以及同一器官在不同患者间形状与大小差异大。本文提出一种新型预训练-微调框架MOSMOS,通过医学报告监督实现多器官分割。首先,在预训练阶段引入全局对比学习,最大化对齐影像与报告的配对;为缓解粒度差异,进一步采用多标签识别隐式学习像素与器官标签间的语义对应关系;更重要的是,预训练模型可通过引入像素-标签注意力图迁移至任意分割模型。我们使用2D U-Net和3D UNETR验证了方法的通用性,并在BTCV、AMOS、MMWHS、BRATS等数据集上,针对多种疾病与模态进行了广泛评估。实验结果表明该框架有效,可作为未来基于医学报告自动标注任务的基础。

原文摘要 · Abstract (English)

Owing to a large amount of multi-modal data in modern medical systems, such as medical images and reports, Medical Vision-Language Pre-training (Med-VLP) has demonstrated incredible achievements in coarse-grained downstream tasks (i.e., medical classification, retrieval, and visual question answering). However, the problem of transferring knowledge learned from Med-VLP to fine-grained multi-organ segmentation tasks has barely been investigated. Multi-organ segmentation is challenging mainly due to the lack of large-scale fully annotated datasets and the wide variation in the shape and size of the same organ between individuals with different diseases. In this paper, we propose a novel pre-training & fine-tuning framework for Multi-Organ Segmentation by harnessing Medical repOrt Supervision (MOSMOS). Specifically, we first introduce global contrastive learning to maximally align the medical image-report pairs in the pre-training stage. To remedy the granularity discrepancy, we further leverage multi-label recognition to implicitly learn the semantic correspondence between image pixels and organ tags. More importantly, our pre-trained models can be transferred to any segmentation model by introducing the pixel-tag attention maps. Different network settings, i.e., 2D U-Net and 3D UNETR, are utilized to validate the generalization. We have extensively evaluated our approach using different diseases and modalities on BTCV, AMOS, MMWHS, and BRATS datasets. Experimental results in various settings demonstrate the effectiveness of our framework. This framework can serve as the foundation to facilitate future research on automatic annotation tasks under the supervision of medical reports.

多器官分割视觉语言模型医学报告预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。