医学图像分割新方法,仅需一张图就能精准生成提示。
Med-PerSAM: One-Shot Visual Prompt Tuning for Personalized Segment Anything Model in Medical Domain
- 基于轻量级形变模型自动优化视觉提示,无需额外训练
- 在多个2D医学影像数据集上超越基础模型和已有SAM方法
- 适合无医学背景者使用,解决提示生成不准与点聚集问题
利用预训练模型结合定制提示进行上下文学习在自然语言处理中已证明高效。受此启发,近期研究将类似方法应用于分割任意模型(SAM)的“单样本”框架,仅需一张参考图像及其标签。然而,这些方法在医学领域面临挑战,主要源于SAM对视觉提示的依赖以及过度依赖像素相似性生成提示,导致(1)提示生成不准确,(2)点提示聚集,影响最终效果。为此,我们提出 extbf{Med-PerSAM},一种专为医学领域设计的简单高效的单样本框架。该方法仅通过视觉提示工程实现,无需微调预训练SAM或人工干预,得益于我们创新的自动化提示生成流程。通过将轻量级形变提示调优模型与SAM结合,实现提示的提取与迭代优化,显著提升预训练SAM性能。该技术在医学领域尤为重要,因非专业人员难以生成有效提示。实验表明,我们的模型在多个2D医学影像数据集上优于多种基础模型及先前的SAM相关方法。
原文摘要 · Abstract (English)
Leveraging pre-trained models with tailored prompts for in-context learning has proven highly effective in NLP tasks. Building on this success, recent studies have applied a similar approach to the Segment Anything Model (SAM) within a ``one-shot" framework, where only a single reference image and its label are employed. However, these methods face limitations in the medical domain, primarily due to SAM's essential requirement for visual prompts and the over-reliance on pixel similarity for generating them. This dependency may lead to (1) inaccurate prompt generation and (2) clustering of point prompts, resulting in suboptimal outcomes. To address these challenges, we introduce \textbf{Med-PerSAM}, a novel and straightforward one-shot framework designed for the medical domain. Med-PerSAM uses only visual prompt engineering and eliminates the need for additional training of the pretrained SAM or human intervention, owing to our novel automated prompt generation process. By integrating our lightweight warping-based prompt tuning model with SAM, we enable the extraction and iterative refinement of visual prompts, enhancing the performance of the pre-trained SAM. This advancement is particularly meaningful in the medical domain, where creating visual prompts poses notable challenges for individuals lacking medical expertise. Our model outperforms various foundational models and previous SAM-based approaches across diverse 2D medical imaging datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。