arXiv:2608.01663cs.CVcs.AI2026-08

用少量图像-掩码对训练视觉提示,提升医学影像分割模型性能

Few-Shot Concept Prompt Learning for Segmentation Foundation Models via Visual Grounding

论文配图:Few-Shot Concept Prompt Learning for Segmentation Foundation Models via Visual Grounding
图 1 · 摘自论文原文
  • 从少量图像-掩码对中学习连续视觉概念提示,冻结主干网络
  • 在4个医学影像数据集上,最高提升Dice分数0.62
  • 无需额外文本数据或重训练,适用于各类基础模型

可提示的分割基础模型(如SAM3和Medical SAM3)通过自然语言接口实现少样本、交互式医学影像分割,但在临床任务中的表现远未达预期。我们认为这一差距并非源于医疗领域预训练不足或提示词表达不精准,而是由于配对图像-文本监督数据稀缺所导致的结构性瓶颈,这在多数临床模态中普遍存在。我们进一步假设,该瓶颈源于以自然语言作为控制信号:若直接从目标分布中学习视觉上对齐的提示,即可恢复性能,且无需额外图像-文本数据或主干网络重训练。为此提出少样本概念提示学习(FS-CPL),通过小规模支持集(K对图像-掩码)与掩码监督,学习连续概念提示嵌入$\mathbf{p}^* \ in \mathbb{R}^{T \ times d}$,主干网络保持冻结。在涵盖超声与内窥镜的四个公开基准(BUSI、HC18、TN3K、CVC-Clinic)上,FS-CPL相比标准文本提示实现最高+0.62的绝对Dice提升,且具备主干网络无关性:同时提升通用SAM3与领域特定预训练的Medical SAM3,表明视觉概念提示与领域预训练具有互补性。

原文摘要 · Abstract (English)

Promptable segmentation foundation models (FMs) such as SAM3 and Medical SAM3 promise few-shot, interactively-specified segmentation for medical imaging through a natural language interface, yet their performance on clinical tasks falls well short of this promise. We posit that this shortfall is not an artefact of insufficient medical pretraining or imperfect prompt phrasing, but a structural limitation that will persist in any domain where paired image-text supervision is scarce, as it is across most clinical modalities. We further hypothesize that the limitation is specific to natural language as a control signal: a visually grounded prompt, learned directly from the target distribution, should recover the lost performance without additional image-text data or backbone retraining. We propose Few-Shot Concept Prompt Learning (FS-CPL), which learns a continuous concept prompt embedding $\mathbf{p}^* \in \mathbb{R}^{T \times d}$ from a small support set of $K$ image--mask pairs via mask supervision, with the encoder-decoder backbone frozen. Across four public benchmarks spanning ultrasound and endoscopy (BUSI, HC18, TN3K, CVC-Clinic), FS-CPL delivers absolute Dice improvements of up to $+0.62$ over canonical text prompts and is \emph{backbone-agnostic}: it lifts both vanilla SAM3 and the domain-specifically pretrained Medical SAM3, showing that visual concept prompting is complementary to in-domain pretraining.

医学影像少样本学习提示学习分割模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。