用双语义监督提升医学图像分割伪标签质量,减少错误标注影响。
DuSSS: Dual Semantic Similarity-Supervised Vision-Language Model for Semi-Supervised Medical Image Segmentation
- 设计双对比学习框架,增强视觉与文本模态间语义一致性。
- 在三个数据集上达到82.52%、74.61%和78.03%的Dice分数。
- 适合需要降低人工标注成本的医疗图像分割研究者。
半监督医学图像分割(SSMIS)通过一致性学习缓解像素级人工标注负担,但常受低质量伪标签错误监督影响。视觉语言模型(VLM)可通过文本提示引入多模态监督信息以增强伪标签,却面临跨模态歧义问题:生成信息可能对应多个目标。为此,本文提出双语义相似性监督视觉语言模型(DuSSS)。首先,设计双对比学习(DCL),通过捕捉各模态内在表征及跨模态语义关联,提升跨模态一致性。其次,提出语义相似性监督策略(SSS),在每个对比学习过程中注入基于不确定性分布的语义相似性监督,鼓励学习多重语义对应关系。此外,构建基于VLM的新型半监督分割网络,利用预训练VLM生成文本引导的监督信息,优化伪标签以实现更优的一致性正则化。实验表明,所提方法在三个公开数据集(QaTa-COV19、BM-Seg、MoNuSeg)上分别取得82.52%、74.61%和78.03%的Dice分数,表现优异。
原文摘要 · Abstract (English)
Semi-supervised medical image segmentation (SSMIS) uses consistency learning to regularize model training, which alleviates the burden of pixel-wise manual annotations. However, it often suffers from error supervision from low-quality pseudo labels. Vision-Language Model (VLM) has great potential to enhance pseudo labels by introducing text prompt guided multimodal supervision information. It nevertheless faces the cross-modal problem: the obtained messages tend to correspond to multiple targets. To address aforementioned problems, we propose a Dual Semantic Similarity-Supervised VLM (DuSSS) for SSMIS. Specifically, 1) a Dual Contrastive Learning (DCL) is designed to improve cross-modal semantic consistency by capturing intrinsic representations within each modality and semantic correlations across modalities. 2) To encourage the learning of multiple semantic correspondences, a Semantic Similarity-Supervision strategy (SSS) is proposed and injected into each contrastive learning process in DCL, supervising semantic similarity via the distribution-based uncertainty levels. Furthermore, a novel VLM-based SSMIS network is designed to compensate for the quality deficiencies of pseudo-labels. It utilizes the pretrained VLM to generate text prompt guided supervision information, refining the pseudo label for better consistency regularization. Experimental results demonstrate that our DuSSS achieves outstanding performance with Dice of 82.52%, 74.61% and 78.03% on three public datasets (QaTa-COV19, BM-Seg and MoNuSeg).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。