arXiv:2511.19759cs.CV2025-11被引 1

用视觉语言模型提升医学图像分割的少样本表现

Vision-Language Enhanced Foundation Model for Semi-Supervised Medical Image Segmentation

  • 引入视觉语言助手,通过模板匹配生成结构化提示
  • 在极少量标注下,分割准确率超越现有最优方法
  • 适合医疗图像少样本分割研究者参考

半监督学习(SSL)已成为降低医学图像分割对专家标注依赖的有效范式。视觉语言模型(VLM)在多种视觉领域展现出强大的泛化和少样本能力。本文将VLM融入半监督医学图像分割框架,提出视觉语言增强的半监督分割助手VESSA,将基础级视觉-语义理解引入SSL。该方法分为两阶段:第一阶段,基于包含金标准样例的模板库,训练具备参考引导能力的分割基础模型VESSA;输入模板对后,其通过视觉特征匹配提取样例分割中的语义与空间线索,生成结构化提示,驱动类SAM掩码解码器生成分割结果。第二阶段,将VESSA作为即插即用的教师模型集成进SSL框架,提供模板引导的伪标签,增强任务特定学生模型的监督信号。在多个数据集与领域上的实验表明,融合VESSA的SSL显著提升分割精度,在极低标注条件下优于当前最佳基线。

原文摘要 · Abstract (English)

Semi-supervised learning (SSL) has emerged as an efficient paradigm for medical image segmentation, reducing the reliance on extensive expert annotations. Vision-language models (VLMs) have demonstrated strong generalization and few-shot capabilities across diverse visual domains. In this work, we integrate a VLM into a semi-supervised medical image segmentation model by adding a Vision-Language Enhanced Semi-supervised Segmentation Assistant (VESSA) that incorporates foundation-level visual-semantic understanding into SSL frameworks. Our approach consists of two stages. In Stage 1, the VLM-enhanced segmentation foundation model VESSA is trained as a reference-guided segmentation assistant using a template bank containing gold-standard exemplars, simulating learning from limited labeled data. Given an input-template pair, VESSA performs visual feature matching to extract representative semantic and spatial cues from exemplar segmentations, generating structured prompts for a Segment Anything Model (SAM)-inspired mask decoder to produce segmentation masks. In Stage 2, VESSA is integrated into an SSL framework as a plug-and-play teacher, providing template-guided pseudo-labels that complement the task-specific student model and strengthen supervision under scarce annotations. Extensive experiments across multiple segmentation datasets and domains show that VESSA-augmented SSL significantly enhances segmentation accuracy, outperforming state-of-the-art baselines under extremely limited annotation conditions.

医学图像半监督视觉语言分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。