arXiv:2604.10000cs.CV2026-04

用文本提示增强医学图像分割,提升模糊区域识别精度

SwinTextUNet: Integrating CLIP-Based Text Guidance into Swin Transformer U-Nets for Medical Image Segmentation

论文配图:SwinTextUNet: Integrating CLIP-Based Text Guidance into Swin Transformer U-Nets for Medical Image Segmentation
图 1 · 摘自论文原文
  • 将CLIP文本嵌入Swin Transformer U-Net,通过跨注意力融合语义文本与视觉特征
  • 在QaTaCOV19数据集上达到86.47%的Dice分数和78.2%的IoU
  • 适合需要高精度分割的临床场景,尤其对低对比度病灶有效

精准的医学图像分割对于辅助诊断和治疗规划至关重要。仅依赖视觉特征的传统模型在面对模糊或低对比度图像时表现不佳。为此,我们提出SwinTextUNet,一种融合对比语言图像预训练(CLIP)文本嵌入的多模态分割框架,集成于Swin Transformer U-Net主干网络中。通过跨注意力机制与卷积融合,模型有效对齐语义文本引导与分层视觉表征,提升了鲁棒性与准确性。我们在QaTaCOV19数据集上评估该方法,其四阶段变体在性能与复杂度间取得最优平衡,获得86.47%的Dice分数与78.2%的IoU。消融实验进一步验证了文本引导与多模态融合的重要性。这些结果凸显了视觉-语言融合在推进医学图像分割及支持临床诊断工具方面的潜力。

原文摘要 · Abstract (English)

Precise medical image segmentation is fundamental for enabling computer aided diagnosis and effective treatment planning. Traditional models that rely solely on visual features often struggle when confronted with ambiguous or low contrast patterns. To overcome these limitations, we introduce SwinTextUNet, a multimodal segmentation framework that incorporates Contrastive Language Image Pretraining (CLIP), derived textual embeddings into a Swin Transformer UNet backbone. By integrating cross attention and convolutional fusion, the model effectively aligns semantic text guidance with hierarchical visual representations, enhancing robustness and accuracy. We evaluate our approach on the QaTaCOV19 dataset, where the proposed four stage variant achieves an optimal balance between performance and complexity, yielding Dice and IoU scores of 86.47% and 78.2%, respectively. Ablation studies further validate the importance of text guidance and multimodal fusion. These findings underscore the promise of vision language integration in advancing medical image segmentation and supporting clinically meaningful diagnostic tools.

医学图像分割多模态CLIPSwin Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。