无需训练,提升手术视频分割精度与抗遮挡能力
Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation
- 引入上下文感知记忆机制,增强对复杂运动的追踪能力
- 在EndoVis2017/2018上分别提升4.36%/6.1%性能
- 适合多器械、长时序手术视频的实时分割场景
手术视频分割是计算机辅助手术中的关键任务,对提升手术质量与患者预后至关重要。近期,图像与视频分割性能显著提升的Segment Anything Model 2(SAM2)框架,在面对手术视频特有的快速器械运动、频繁遮挡及复杂器械-组织交互时,其贪婪选择的记忆设计局限被放大,导致长视频分割性能下降。为此,我们提出无需训练的Memory Augmented (MA)-SAM2策略,引入新型上下文感知与遮挡鲁棒的记忆模型。MA-SAM2在复杂器械运动下保持高精度,有效应对遮挡与交互。采用多目标单循环单提示推理,进一步提升多器械视频的追踪效率。不增加参数且无需额外训练,在EndoVis2017和EndoVis2018数据集上分别实现4.36%和6.1%的性能提升,展现出实际手术应用潜力。
原文摘要 · Abstract (English)
Surgical video segmentation is a critical task in computer-assisted surgery, essential for enhancing surgical quality and patient outcomes. Recently, the Segment Anything Model 2 (SAM2) framework has demonstrated remarkable advancements in both image and video segmentation. However, the inherent limitations of SAM2's greedy selection memory design are amplified by the unique properties of surgical videos-rapid instrument movement, frequent occlusion, and complex instrument-tissue interaction-resulting in diminished performance in the segmentation of complex, long videos. To address these challenges, we introduce Memory Augmented (MA)-SAM2, a training-free video object segmentation strategy, featuring novel context-aware and occlusion-resilient memory models. MA-SAM2 exhibits strong robustness against occlusions and interactions arising from complex instrument movements while maintaining accuracy in segmenting objects throughout videos. Employing a multi-target, single-loop, one-prompt inference further enhances the efficiency of the tracking process in multi-instrument videos. Without introducing any additional parameters or requiring further training, MA-SAM2 achieved performance improvements of 4.36% and 6.1% over SAM2 on the EndoVis2017 and EndoVis2018 datasets, respectively, demonstrating its potential for practical surgical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。