利用解剖结构跨个体空间一致性提升3D多模态医学图像自监督学习
Beyond Instance-Level Self-Supervision in 3D Multi-Modal Medical Imaging

- 通过跨实例解剖拓扑一致性构建自监督信号,避免传统方法忽略空间关系的缺陷
- 在7个下游任务中,分割与分类准确率分别提升1.1%和5.94%
- 在测试时模态缺失情况下仍表现鲁棒,适合临床实际应用
医学图像自监督预训练方法通常将每个个体视为独立样本,通过数据增强或掩码重建学习表征。然而,它们未能充分利用生理特征的关键特性:解剖结构在个体间保持一致的空间关系(如丘脑始终位于基底节内侧),无论脑部大小、形状或病理变化如何。本文提出利用这种跨实例拓扑一致性作为监督信号。由于医学影像存在显著的个体与模态差异,我们设计两种对齐机制:(i) 个体内部:在像素级对应可用时,采用跨模态三元组损失显式保留局部邻域拓扑;(ii) 个体之间:无标注对应时,通过伪对应控制部分邻域对齐,防止跨模态拓扑坍缩。我们在7个下游多模态任务上验证方法,分割任务平均提升1.1%,分类任务提升5.94%,且在测试时模态缺失情况下表现出显著更强的鲁棒性。
原文摘要 · Abstract (English)
Self-supervised pre-training methods in medical imaging typically treat each individual as an isolated instance, learning representations through augmentation-based objectives or masked reconstruction. They often do not adequately capitalize on a key characteristic of physiological features: anatomical structures maintain consistent spatial relationships across individuals (instances), such as the thalamus being medial to the basal ganglia, regardless of variations in brain size, shape, or pathology. We propose leveraging this cross-instance topological consistency as a supervisory signal. The challenge arises from the inherent variability in medical imaging, which can differ significantly across instances and modalities. To tackle this, we focus on two alignment regimes. (i) Intra-instance: with pixel-level correspondences available, a cross-modal triplet objective explicitly preserves local neighborhood topology. (ii) Inter-instance: without such supervision, we derive pseudo-correspondences to control partial neighborhood alignment and prevent topology collapse across modalities. We validate our approach across 7 downstream multi-modal tasks, achieving average improvements of 1.1% and 5.94% in segmentation and classification tasks, respectively, and demonstrating significantly better robustness when modalities are missing at test time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。