用文本引导跨模态医学图像分割,提升无监督域适应效果
TCSA-UDA: Text-Driven Cross-Semantic Alignment for Unsupervised Domain Adaptation in Medical Image Segmentation
- 通过文本提示引导视觉特征学习,对齐图像与语义关系
- 在心脏、腹部和脑肿瘤数据集上平均性能提升6.2%以上
- 适合需要跨模态医疗影像分析的研究者和临床应用
医学图像分割中的无监督域适应(UDA)因不同成像模态(如CT与MRI)间存在显著域偏移而面临挑战。尽管近期视觉-语言表示学习方法在医学图像分析中展现潜力,其在跨模态UDA中的作用仍待深入探索。为此,本文提出TCSA-UDA:一种文本驱动的跨语义对齐框架,利用模态感知的文本提示指导域不变视觉表征学习。具体而言,引入视觉-语言协方差余弦损失(VLCoL),将类间视觉特征关系与文本导出的语义关系对齐,促使图像编码器学习具有语义结构且模态鲁棒的表示。此外,设计原型对齐模块,通过高阶类别原型对齐减少源域与目标域间的残余类别差异。在跨模态心脏、腹部及脑肿瘤分割基准上的大量实验表明,TCSA-UDA持续提升适应性能,优于当前最优的UDA方法。结果验证了语言驱动语义引导在域适应医学图像分割中的潜力。代码已开源。
原文摘要 · Abstract (English)
Unsupervised domain adaptation (UDA) for medical image segmentation remains challenging due to substantial domain shifts across imaging modalities, such as CT and MRI. Although recent vision-language representation learning methods have shown promise in medical image analysis, their role in cross-modality UDA segmentation remains underexplored. To address this problem, we propose TCSA-UDA, a Text-driven Cross-Semantic Alignment framework that uses modality-aware textual prompting to guide domain-invariant visual representation learning. Specifically, we introduce a vision-language covariance cosine loss (VLCoL) that aligns inter-class visual feature relationships with text-derived semantic relationships, encouraging the image encoder to learn semantically structured and modality-robust representations. In addition, we incorporate a prototype alignment module to reduce residual class-level discrepancies between source and target domains by aligning high-level class prototypes. Extensive experiments on cross-modality cardiac, abdominal, and brain tumor segmentation benchmarks demonstrate that TCSA-UDA consistently improves adaptation performance and outperforms state-of-the-art UDA methods. These results highlight the potential of language-driven semantic guidance for domain-adaptive medical image segmentation. The code is available at https://github.com/lalitmaurya47/TCSA_UDA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。