arXiv:2501.00525cs.CV2025-01被引 4

首次评估SAM2在9个手术数据集上的零样本分割能力

Systematic Evaluation and Guidelines for Segment Anything Model in Surgical Video Analysis

  • 用点、框、掩码提示与稠密/稀疏微调策略测试SAM2
  • 在器械和多器官分割中表现良好,但动态条件下效果下降
  • 适合关注手术视频分析的AI研究者和临床应用开发者

手术视频分割对AI理解手术中的时空动态至关重要,但受限于标注数据稀缺。基于自然视频预训练的SAM2模型具备零样本手术分割潜力,但在组织变形和器械多样性等复杂手术环境下的适用性尚未探索。我们首次系统评估了SAM2在9个手术数据集(17种手术类型)中的零样本能力,涵盖腹腔镜、内窥镜和机器人手术。分析了点、框、掩码等多种提示方式及稠密、稀疏微调策略,考察其对手术挑战的鲁棒性以及跨手术类型和解剖结构的泛化能力。结果表明,尽管SAM2在结构化场景(如器械分割、多器官分割、场景分割)中表现出显著零样本适应性,但在动态手术条件下性能波动,暴露出对时序一致性与领域特定伪影处理的不足。这些发现为手术数据科学中自适应、数据高效解决方案的发展指明了方向。

原文摘要 · Abstract (English)

Surgical video segmentation is critical for AI to interpret spatial-temporal dynamics in surgery, yet model performance is constrained by limited annotated data. The SAM2 model, pretrained on natural videos, offers potential for zero-shot surgical segmentation, but its applicability in complex surgical environments, with challenges like tissue deformation and instrument variability, remains unexplored. We present the first comprehensive evaluation of the zero-shot capability of SAM2 in 9 surgical datasets (17 surgery types), covering laparoscopic, endoscopic, and robotic procedures. We analyze various prompting (points, boxes, mask) and {finetuning (dense, sparse) strategies}, robustness to surgical challenges, and generalization across procedures and anatomies. Key findings reveal that while SAM2 demonstrates notable zero-shot adaptability in structured scenarios (e.g., instrument segmentation, {multi-organ segmentation}, and scene segmentation), its performance varies under dynamic surgical conditions, highlighting gaps in handling temporal coherence and domain-specific artifacts. These results highlight future pathways to adaptive data-efficient solutions for the surgical data science field.

手术分割零样本SAM2视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。