arXiv:2503.03942cs.CV2025-03被引 3

微调SAM2模型,实现手术视频中器官组织的精准分割与检测。

SurgiSAM2: Fine-tuning a foundational model for surgical video anatomy segmentation and detection

  • 基于5个公开数据集微调SAM2的图像编码器和掩码解码器。
  • 在400样本/类下使用10个提示点,达到0.92的验证集Dice系数。
  • 对未见器官类有强泛化能力,80%类别超越现有最先进方法。

我们评估了SAM2在手术场景理解中的表现,重点考察其在零样本和微调后对器官/组织的语义分割能力。利用五个公开数据集对SAM2进行微调,仅用每类50至400样本,训练受限于真实数据获取条件。通过加权平均骰子系数(WMDC)评估数据量对性能的影响,并与此前报道的最先进(SOTA)结果对比。结果表明,微调后的SurgiSAM2相比基线SAM2相对提升17.9%的WMDC;当提示点数从1增至10、训练样本量从50/类增至400/类时性能提升显著,最佳验证集WMDC达0.92。在测试集上,该模型在30类中的24类(80%)优于先前SOTA,WMDC为0.91。尤其在未见过的器官类别上表现优异,9类中有7类(77.8%)达到最先进水平。结论显示,SAM2在手术场景分割任务中展现出卓越的零样本与微调性能,可推动自动化/半自动化标注流程,降低标注负担,助力多种外科应用。

原文摘要 · Abstract (English)

Background: We evaluate SAM 2 for surgical scene understanding by examining its semantic segmentation capabilities for organs/tissues both in zero-shot scenarios and after fine-tuning. Methods: We utilized five public datasets to evaluate and fine-tune SAM 2 for segmenting anatomical tissues in surgical videos/images. Fine-tuning was applied to the image encoder and mask decoder. We limited training subsets from 50 to 400 samples per class to better model real-world constraints with data acquisition. The impact of dataset size on fine-tuning performance was evaluated with weighted mean Dice coefficient (WMDC), and the results were also compared against previously reported state-of-the-art (SOTA) results. Results: SurgiSAM 2, a fine-tuned SAM 2 model, demonstrated significant improvements in segmentation performance, achieving a 17.9% relative WMDC gain compared to the baseline SAM 2. Increasing prompt points from 1 to 10 and training data scale from 50/class to 400/class enhanced performance; the best WMDC of 0.92 on the validation subset was achieved with 10 prompt points and 400 samples per class. On the test subset, this model outperformed prior SOTA methods in 24/30 (80%) of the classes with a WMDC of 0.91 using 10-point prompts. Notably, SurgiSAM 2 generalized effectively to unseen organ classes, achieving SOTA on 7/9 (77.8%) of them. Conclusion: SAM 2 achieves remarkable zero-shot and fine-tuned performance for surgical scene segmentation, surpassing prior SOTA models across several organ classes of diverse datasets. This suggests immense potential for enabling automated/semi-automated annotation pipelines, thereby decreasing the burden of annotations facilitating several surgical applications.

手术分割SAM2微调医学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。