arXiv:2606.15802cs.CV2026-06

用文字提示提升脊柱分割伪标签质量,仅用5%标注数据达80.44%分割精度

CPS4: Class Prompt driven Semi-Supervised Spine Segmentation with Class-specific Consistency Constraint

论文配图:CPS4: Class Prompt driven Semi-Supervised Spine Segmentation with Class-specific Consistency Constraint
图 1 · 摘自论文原文
  • 通过文本提示与脊柱区域的注意力一致性约束,增强语义对齐
  • 在仅5%标注数据下,脊柱分割Dice达80.44%,超越现有方法
  • 适合需要少样本高质量医学图像分割的研究者使用

视觉语言模型(VLM)有潜力通过文本类别提示生成分割图来提升半监督脊柱分割的伪标签质量,但尚未有人研究。尽管前景可观,但缺乏显式约束以保证类别提示与脊柱结构之间的一致性,导致多类别分割效果不佳。本文提出CPS4,首个基于文本提示的半监督脊柱分割网络,通过类别提示优化伪标签质量。具体包含两个训练阶段:(i) 类别特定一致性约束的VLM预训练阶段:提出词元与像素级注意力损失,优化类别提示与脊柱单元在语义空间的一致性;(ii) 文本提示驱动的半监督分割阶段:利用预训练的视觉-文本编码器,为未标注脊柱图像生成各分类二值分割图,并融合为统一多类别分割图,显著提升伪标签质量。实验表明,仅使用5%标注数据,CPS4在公开脊柱分割数据集上取得80.44%的Dice分数,优于主流半监督学习与VLM方法。代码将开源。

原文摘要 · Abstract (English)

Vision Language Model (VLM) has great potential to enhance the quality of pseudo labels in semi-supervised spine segmentation by leveraging textual class prompts to generate segmentation map, but no one has studied it yet. Although promising, it lacks explicit constraints to ensure consistency between spine class prompts and spine unit region, resulting in unsatisfactory performance in multi-class segmentation map generation. In this paper, we propose CPS4, the first text-guided semi-supervised spine segmentation network using class prompts to enhance the quality of spine pseudo labels. Specifically, CPS4 is implemented through two training stages. (i) Class-specific consistency constrained VLM pretraining stage: we propose token- and pixel-level attention loss to optimize the consistency between class prompts and spine units, forcing the textual class prompt to be closely coupled with the target spine unit in the semantic space. (ii) Class Prompt driven semi-supervised spine segmentation stage: using the pretrained vision-text encoder, we derive each class-specific binary segmentation map for the unlabeled spine image and integrate them into an unified multi-class segmentation map, improving the quality of the spine pseudo label generated by the semi-supervised spine segmentation network. Experimental results show that our CPS4 achieves superior spine segmentation performance with Dice of 80.44%, only using 5% labeled data on the public spine segmentation dataset, surpassing popular semi-supervised learning and VLM methods. Our code will be available.

医学图像分割半监督学习视觉语言模型少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。