arXiv:2507.12382cs.CV2025-07中稿 · MICCAI 2025被引 5

用文本增强3D医学图像分割,提升小样本下的精度

Text-driven Multiplanar Visual Interaction for Semi-supervised Medical Image Segmentation

  • 通过多平面映射融合文本与视觉信息,增强特征类别感知
  • 在三个公开数据集上显著优于现有方法,性能提升明显
  • 适合医疗图像标注稀缺场景,尤其对半监督学习研究者有用

半监督医学图像分割能有效缓解数据标注成本高的问题。当标注数据有限时,文本信息可提供额外语义上下文以增强视觉理解。然而,利用文本信息提升3D医学图像任务中视觉嵌入的研究仍较少。本文提出一种新型文本驱动的多平面视觉交互框架(Text-SemiSeg),包含三个模块:文本增强多平面表征(TMR)、类别感知语义对齐(CSA)和动态认知增强(DCA)。TMR通过平面映射实现文本-视觉交互,增强视觉特征的类别敏感性;CSA引入可学习变量,对齐文本特征与视觉中间层特征;DCA通过标签与无标签数据交互,减少分布差异,提升模型鲁棒性。在三个公开数据集上的实验表明,该模型能有效利用文本信息增强视觉特征,性能优于现有方法。代码已开源:https://github.com/taozh2017/Text-SemiSeg。

原文摘要 · Abstract (English)

Semi-supervised medical image segmentation is a crucial technique for alleviating the high cost of data annotation. When labeled data is limited, textual information can provide additional context to enhance visual semantic understanding. However, research exploring the use of textual data to enhance visual semantic embeddings in 3D medical imaging tasks remains scarce. In this paper, we propose a novel text-driven multiplanar visual interaction framework for semi-supervised medical image segmentation (termed Text-SemiSeg), which consists of three main modules: Text-enhanced Multiplanar Representation (TMR), Category-aware Semantic Alignment (CSA), and Dynamic Cognitive Augmentation (DCA). Specifically, TMR facilitates text-visual interaction through planar mapping, thereby enhancing the category awareness of visual features. CSA performs cross-modal semantic alignment between the text features with introduced learnable variables and the intermediate layer of visual features. DCA reduces the distribution discrepancy between labeled and unlabeled data through their interaction, thus improving the model's robustness. Finally, experiments on three public datasets demonstrate that our model effectively enhances visual features with textual information and outperforms other methods. Our code is available at https://github.com/taozh2017/Text-SemiSeg.

医学图像半监督文本增强3D分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。