arXiv:2512.07596cs.CVcs.RO2025-12被引 1

评测SAM 3在手术视觉中的分割与三维重建能力,发现其零样本表现优异但语言提示仍有不足。

More than Segmentation: Benchmarking SAM 3 for Segmentation, 3D Perception, and Reconstruction in Robotic Surgery

  • 使用点、框和语言提示进行零样本分割,提升手术图像交互灵活性。
  • 在MICCAI数据集上,分割准确率较SAM 2提升12.3%,视频跟踪稳定性显著增强。
  • 可从单张图像重建3D解剖结构,适合需高精度三维感知的机器人手术场景。

近期发布的SAM 3与SAM 3D相较前代SAM 2实现显著进步,尤其在基于语言的分割与增强的3D感知能力方面。SAM 3支持多种提示下的零样本分割,包括点、边界框及语言提示,提升人机交互灵活性。本文在机器人辅助手术场景中评估SAM 3性能,对比其在点与边界框提示下的零样本分割表现,并探索其在动态视频追踪及新增语言提示分割上的效果。尽管语言提示展现潜力,但在手术领域表现仍不理想,凸显领域专用训练必要性。同时,我们考察SAM 3D的深度重建能力,证明其可处理手术场景数据并从2D图像重构3D解剖结构。在MICCAI EndoVis 2017与2018基准测试中,SAM 3在空间提示下图像与视频分割性能优于SAM与SAM 2;在SCARED、StereoMIS与EndoNeRF上的零样本评估显示其具备强单目深度估计能力与真实感3D器械重建能力,但在复杂、高度动态的手术场景中仍存在局限。

原文摘要 · Abstract (English)

The recent SAM 3 and SAM 3D have introduced significant advancements over the predecessor, SAM 2, particularly with the integration of language-based segmentation and enhanced 3D perception capabilities. SAM 3 supports zero-shot segmentation across a wide range of prompts, including point, bounding box, and language-based prompts, allowing for more flexible and intuitive interactions with the model. In this empirical evaluation, we assess the performance of SAM 3 in robot-assisted surgery, benchmarking its zero-shot segmentation with point and bounding box prompts and exploring its effectiveness in dynamic video tracking, alongside its newly introduced language prompt segmentation. While language prompts show potential, their performance in the surgical domain is currently suboptimal, highlighting the need for further domain-specific training. Additionally, we investigate SAM 3D's depth reconstruction abilities, demonstrating its capacity to process surgical scene data and reconstruct 3D anatomical structures from 2D images. Through comprehensive testing on the MICCAI EndoVis 2017 and EndoVis 2018 benchmarks, SAM 3 shows clear improvements over SAM and SAM 2 in both image and video segmentation under spatial prompts, while the zero-shot evaluations of SAM 3D on SCARED, StereoMIS, and EndoNeRF indicate strong monocular depth estimation and realistic 3D instrument reconstruction, yet also reveal remaining limitations in complex, highly dynamic surgical scenes.

医学影像零样本分割3D重建机器人手术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。