arXiv:2506.18042cs.CV2025-06

用稀疏标注实现精准医学图像分割,提升小病灶识别能力

CmFNet: Cross-modal Fusion Network for Weakly-supervised Segmentation of Medical Images

  • 通过跨模态融合提取多模态图像互补特征
  • 在鼻咽癌和腹部器官数据集上显著优于现有弱监督方法
  • 适合临床诊疗中缺乏密集标注的场景

精确的自动医学图像分割依赖高质量、密集的标注,但其成本高、耗时长。弱监督学习通过使用稀疏、粗略的标注提供更高效替代方案,但仍面临性能下降与过拟合问题。为此,我们提出CmFNet,一种新型3D弱监督跨模态医学图像分割方法。CmFNet包含三个核心组件:模态特异性特征学习网络、跨模态特征学习网络和混合监督策略。前者与后者协同整合多模态图像的互补信息,增强跨模态共享特征,从而提升分割性能。混合监督策略通过笔画监督、模态内正则化与模态间一致性,建模空间与上下文关系并促进特征对齐,有效缓解过拟合,实现鲁棒分割。该方法在临床交叉模态鼻咽癌(含CT与MR)数据集及公开的CT全腹器官数据集(WORD)上均优于当前最先进弱监督方法。此外,在使用完整标注时,其表现甚至超过全监督方法。该方法可助力临床治疗,惠及物理师、放射科医生、病理科医生与肿瘤科医生。

原文摘要 · Abstract (English)

Accurate automatic medical image segmentation relies on high-quality, dense annotations, which are costly and time-consuming. Weakly supervised learning provides a more efficient alternative by leveraging sparse and coarse annotations instead of dense, precise ones. However, segmentation performance degradation and overfitting caused by sparse annotations remain key challenges. To address these issues, we propose CmFNet, a novel 3D weakly supervised cross-modal medical image segmentation approach. CmFNet consists of three main components: a modality-specific feature learning network, a cross-modal feature learning network, and a hybrid-supervised learning strategy. Specifically, the modality-specific feature learning network and the cross-modal feature learning network effectively integrate complementary information from multi-modal images, enhancing shared features across modalities to improve segmentation performance. Additionally, the hybrid-supervised learning strategy guides segmentation through scribble supervision, intra-modal regularization, and inter-modal consistency, modeling spatial and contextual relationships while promoting feature alignment. Our approach effectively mitigates overfitting, delivering robust segmentation results. It excels in segmenting both challenging small tumor regions and common anatomical structures. Extensive experiments on a clinical cross-modal nasopharyngeal carcinoma (NPC) dataset (including CT and MR imaging) and the publicly available CT Whole Abdominal Organ dataset (WORD) show that our approach outperforms state-of-the-art weakly supervised methods. In addition, our approach also outperforms fully supervised methods when full annotation is used. Our approach can facilitate clinical therapy and benefit various specialists, including physicists, radiologists, pathologists, and oncologists.

弱监督医学分割跨模态影像分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。