arXiv:2508.18613eess.IVcs.LG2025-08被引 1

利用影像模态和解剖部位信息提升医学图像预训练效果

ModAn-MulSupCon: Modality-and Anatomy-Aware Multi-Label Supervised Contrastive Pretraining for Medical Imaging

  • 将模态与解剖部位编码为多标签向量,构建对比学习目标
  • 在膝关节MRI、乳腺与甲状腺超声数据上实现0.964~0.763的最优AUC
  • 适合有微调条件的临床场景,尤其标签稀缺时表现优异

专家标注限制了医学影像的大规模监督预训练,而普遍存在的元数据(如模态、解剖区域)仍被低估。本文提出ModAn-MulSupCon,一种面向模态与解剖结构的多标签监督对比预训练方法,利用这些元数据学习可迁移表征。每个图像的模态与解剖区域被编码为多热向量,使用ResNet-18在小规模RadImageNet(miniRIN,16,222张图像)上进行预训练,采用杰卡德加权的多标签监督对比损失。随后在三个二分类任务中评估:前交叉韧带撕裂(膝关节MRI)、病灶恶性程度(乳腺超声)、结节恶性程度(甲状腺超声)。微调结果表明,ModAn-MulSupCon在MRNet-ACL(AUC=0.964)和甲状腺数据集(AUC=0.763)上优于所有基线(p<0.05),在乳腺数据集上排名第二(0.926,仅次于SimCLR的0.940,差异不显著)。当编码器冻结时,SimCLR/ImageNet表现更优,说明ModAn-MulSupCon的表征优势在于任务适配而非线性可分性。结论:将易得的模态/解剖元数据作为多标签目标,提供了一种实用且可扩展的预训练信号,在可微调场景下显著提升下游性能。

原文摘要 · Abstract (English)

Background and objective: Expert annotations limit large-scale supervised pretraining in medical imaging, while ubiquitous metadata (modality, anatomical region) remain underused. We introduce ModAn-MulSupCon, a modality- and anatomy-aware multi-label supervised contrastive pretraining method that leverages such metadata to learn transferable representations. Method: Each image's modality and anatomy are encoded as a multi-hot vector. A ResNet-18 encoder is pretrained on a mini subset of RadImageNet (miniRIN, 16,222 images) with a Jaccard-weighted multi-label supervised contrastive loss, and then evaluated by fine-tuning and linear probing on three binary classification tasks--ACL tear (knee MRI), lesion malignancy (breast ultrasound), and nodule malignancy (thyroid ultrasound). Result: With fine-tuning, ModAn-MulSupCon achieved the best AUC on MRNet-ACL (0.964) and Thyroid (0.763), surpassing all baselines ($p<0.05$), and ranked second on Breast (0.926) behind SimCLR (0.940; not significant). With the encoder frozen, SimCLR/ImageNet were superior, indicating that ModAn-MulSupCon representations benefit most from task adaptation rather than linear separability. Conclusion: Encoding readily available modality/anatomy metadata as multi-label targets provides a practical, scalable pretraining signal that improves downstream accuracy when fine-tuning is feasible. ModAn-MulSupCon is a strong initialization for label-scarce clinical settings, whereas SimCLR/ImageNet remain preferable for frozen-encoder deployments.

医学影像对比学习多标签预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。