用增强影像引导低对比度影像分割,提升肿瘤定位精度。
ViPSAM: Visual Prompting Medical Image Segmentation Using Segment Anything Model

- 引入跨模态视觉提示,融合增强与非增强图像特征
- 在肝部病变分割任务中显著优于U-Net与SAM基线方法
- 适合放射治疗中低对比度CT的精准病灶勾画场景
在质子治疗计划中,呼吸门控非增强CT(NCCT)常用于病灶分割,但由于病灶与背景对比度低,精确勾画仍具挑战。尽管基于学习的方法表现良好,但在非增强图像分割上仍存困难。受临床实践启发——即利用增强MRI辅助在NCCT上勾画病灶——我们提出ViPSAM,一种基于分割一切模型(SAM)的视觉提示框架。该框架引入视觉提示编码器,从增强图像中提取引导特征,并设计视觉引导交叉注意力模块,融合非增强与增强图像特征,从而增强低对比区域的病灶相关表征。掩码解码器亦通过参数高效方式适配,以有效利用视觉提示。我们在质子治疗中获取的肝部病变NCCT数据集上评估该方法。实验结果表明,ViPSAM优于代表性U-Net与SAM基线方法,证明跨模态视觉提示可实现更鲁棒、更准确的非增强图像分割。
原文摘要 · Abstract (English)
In proton therapy planning, respiratory-gated non-contrast CT (NCCT) is commonly used for lesion segmentation; however, accurate delineation remains challenging due to low lesion-to-background contrast. Although learning-based methods have shown strong performance, they often struggle with non-contrast image segmentation. Inspired by clinical practice, where contrast-enhanced MRI is referenced to delineate lesions on NCCT, we propose ViPSAM, a visual prompting framework that leverages complementary cross-modality information. Built upon the Segment Anything Model (SAM), ViPSAM introduces a visual prompt encoder to extract guidance features from contrast-enhanced images and a visual-guided cross-attention module to integrate non-contrast and contrast-enhanced features, thereby enhancing lesion-relevant representations in low-contrast regions. The mask decoder is further adapted in a parameter-efficient manner to utilize visual prompts effectively. We evaluate the proposed method on liver lesion segmentation using NCCT acquired for proton therapy. Experimental results demonstrate that ViPSAM outperforms representative U-Net- and SAM-based methods, indicating that cross-modality visual prompting enables more robust and accurate segmentation in non-contrast images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。