用非目标数据自动生成提示,让SAM在医学分割中更智能、更易交互。
Proxy Prompt: Endowing SAM and SAM 2 with Auto-Interactive-Prompt for Medical Segmentation
- 基于非目标数据自动生成提示,无需人工标注
- 仅用16张标注图即达顶尖性能,媲美全量训练模型
- 通过上下文色彩化增强人机互动,突出用户定义对象
本文旨在解决SAM和SAM2在临床应用中自动化提示与人机交互能力不足的问题。我们提出代理提示(Proxy Prompt, PP),利用带有预标注掩码的非目标数据自动生成提示。设计了一种三步式上下文选择策略,通过视觉马尔可夫网络和选择性图谱,从非目标数据中自适应选取最具代表性的上下文信息,提升非目标图像-掩码对在目标图像/视频上的分割引导能力。为强化人机交互,进一步引入一种双反向交叉注意力的上下文色彩化模块,增强目标特征与上下文嵌入间的交互,并放大用户定义对象的显著特征。在四个公开数据集上的广泛评估表明,本方法达到当前最优性能,且在仅使用16张图像掩码的情况下,结果可媲美全量训练模型。
原文摘要 · Abstract (English)
In this paper, we aim to address the unmet demand for automated prompting and enhanced human-model interactions of SAM and SAM2 for the sake of promoting their widespread clinical adoption. Specifically, we propose Proxy Prompt (PP), auto-generated by leveraging non-target data with a pre-annotated mask. We devise a novel 3-step context-selection strategy for adaptively selecting the most representative contextual information from non-target data via vision mamba and selective maps, empowering the guiding capability of non-target image-mask pairs for segmentation on target image/video data. To reinforce human-model interactions in PP, we further propose a contextual colorization module via a dual-reverse cross-attention to enhance interactions between target features and contextual-embedding with amplifying distinctive features of user-defined object(s). Via extensive evaluations, our method achieves state-of-the-art performance on four public datasets and yields comparable results with fully-trained models, even when trained with only 16 image masks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。