arXiv:2511.21926eess.IVcs.CV2025-11被引 3

对比SAM 2与SAM 3在3D医学图像零样本分割中的表现,发现SAM 3整体更优。

Comparing SAM 2 and SAM 3 for Zero-Shot Segmentation of 3D Medical Data

  • 在多种提示策略下系统评估两个模型的分割性能。
  • SAM 3在点击提示下表现更稳,过分割和预测滞留问题更少。
  • 适用于多数医学影像任务,尤其在超声与内镜序列中优势明显。

基础模型如分割一切模型(SAM)引发了对可提示零样本分割的广泛关注。尽管这些模型在自然图像上表现优异,但在医学数据上的行为仍缺乏充分描述。虽然SAM 2已被广泛用于3D医学工作流标注,但新发布的SAM 3引入了新架构,可能改变视觉提示的解析与传播方式。为评估SAM 3能否作为SAM 2在3D医学数据零样本分割中的即插即用替代品,我们首次在多种提示策略下,以提示性视觉分割(PVS)模式对SAM 3进行系统比较。我们在16个公开数据集(包括CT、MRI、超声、内镜)上进行基准测试,覆盖54个解剖结构、病灶及手术器械。进一步量化了三种失败模式:提示帧过分割、目标消失后的过传播、以及良好初始化预测的持续保留。结果显示,SAM 3在各类模态的点击提示下始终更强,过分割失败更少,预测衰减更慢;在边界框与掩码提示下,性能差距缩小,部分结构中两者在终止行为上出现权衡,但SAM 3在超声与内镜序列中仍保持优势。总体结果表明,SAM 3是大多数医学分割任务的更优默认选择,同时明确了何时仍宜选用SAM 2作为传播器。

原文摘要 · Abstract (English)

Foundation models, such as the Segment Anything Model (SAM), have heightened interest in promptable zero-shot segmentation. Although these models perform strongly on natural images, their behavior on medical data remains insufficiently characterized. While SAM 2 has been widely adopted for annotation in 3D medical workflows, the recently released SAM 3 introduces a new architecture that may change how visual prompts are interpreted and propagated. Therefore, to assess whether SAM 3 can serve as an out-of-the-box replacement for SAM 2 for zero-shot segmentation of 3D medical data, we present the first controlled comparison of both models by evaluating SAM 3 in its Promptable Visual Segmentation (PVS) mode using a variety of prompting strategies. We benchmark on 16 public datasets (CT, MRI, Ultrasound, endoscopy) covering 54 anatomical structures, pathologies, and surgical instruments. We further quantify three failure modes: prompt-frame over-segmentation, over-propagation after object disappearance, and temporal retention of well-initialized predictions. Our results show that SAM 3 is consistently stronger under click prompting across modalities, with fewer prompt-frame over-segmentation failures and slower prediction retention decay compared to SAM 2. Under bounding-box and mask prompts, performance gaps narrow in few structures of CT/MR and the models trade off termination behavior, while SAM 3 remains stronger on ultrasound and endoscopy sequences. The overall results position SAM 3 as the superior default choice for most medical segmentation tasks, while clarifying when SAM 2 remains a preferable propagator.

医学图像零样本分割SAM模型3D分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。