用语言引导自适应融合,少样本实现肺动脉静脉3D分割
Self-adaptive vision-language model for 3D segmentation of pulmonary artery and vein
- 用CLIP模型提取特征,通过自适应适配器融合图文信息
- 在718例数据上达到领先效果,显著减少标注依赖
- 适合医疗图像少样本分割研究者参考
准确分割肺部结构对临床诊断、疾病研究和治疗规划至关重要。尽管基于深度学习的分割技术已取得进展,但多数方法需要大量标注数据。因此,开发少标注数据下的高精度分割方法尤为重要。近年来,预训练视觉-语言基础模型(如CLIP)为通用计算机视觉任务提供了新可能,能在少量标注下实现优异性能。然而,其在肺动脉-静脉分割中的应用仍有限。本文提出一种名为语言引导自适应交叉注意力融合框架的新方法。该方法采用预训练的CLIP作为强特征提取器,对3D CT扫描进行分割,并自适应融合文本与图像表示的跨模态信息。设计特殊适配模块,结合自适应学习策略,有效融合两种模态嵌入。我们在目前最大的肺动脉-静脉CT数据集(共718个标注样本)上进行了充分验证,实验结果表明,本方法显著优于现有最先进方法。数据与代码将在论文接收后公开。
原文摘要 · Abstract (English)
Accurate segmentation of pulmonary structures iscrucial in clinical diagnosis, disease study, and treatment planning. Significant progress has been made in deep learning-based segmentation techniques, but most require much labeled data for training. Consequently, developing precise segmentation methods that demand fewer labeled datasets is paramount in medical image analysis. The emergence of pre-trained vision-language foundation models, such as CLIP, recently opened the door for universal computer vision tasks. Exploiting the generalization ability of these pre-trained foundation models on downstream tasks, such as segmentation, leads to unexpected performance with a relatively small amount of labeled data. However, exploring these models for pulmonary artery-vein segmentation is still limited. This paper proposes a novel framework called Language-guided self-adaptive Cross-Attention Fusion Framework. Our method adopts pre-trained CLIP as a strong feature extractor for generating the segmentation of 3D CT scans, while adaptively aggregating the cross-modality of text and image representations. We propose a s pecially designed adapter module to fine-tune pre-trained CLIP with a self-adaptive learning strategy to effectively fuse the two modalities of embeddings. We extensively validate our method on a local dataset, which is the largest pulmonary artery-vein CT dataset to date and consists of 718 labeled data in total. The experiments show that our method outperformed other state-of-the-art methods by a large margin. Our data and code will be made publicly available upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。