提升手术视频中器械分割的准确性,尤其在遮挡后能准确恢复身份。
ReMeDI: Refined Memory for Disambiguation of Identities with SAM3 in Surgical Segmentation
- 通过感知相关性的记忆筛选和分段插值,扩展有效记忆容量。
- 遮挡后通过特征重识别与时间投票,实现身份准确恢复,零样本下提升约5.8%-8%。
- 无需训练,可直接部署于SAM3,适合临床实时手术辅助系统。
内窥镜手术中精确的器械分割对计算机辅助干预至关重要,但频繁遮挡、快速运动及器械反复进入使任务极具挑战。尽管SAM3提供了强大的时空视频对象分割框架,其在手术场景中的表现受限于无差别记忆更新、固定记忆容量以及遮挡后的身份恢复能力弱。本文提出ReMeDI-SAM3,一种无需训练的SAM3扩展方法,包含三项改进:(i) 基于相关性的记忆过滤与专用遮挡感知记忆,用于存储遮挡前帧;(ii) 分段插值方案,提升有效记忆容量;(iii) 基于特征的重识别模块结合时间投票机制,实现遮挡后可靠的身份判别。三者协同减少误差累积,支持遮挡后稳定恢复。在EndoVis17、EndoVis18和CholecSeg8k数据集的零样本设置下,相较原始SAM3,mciou分别提升约5.8%、8%和2%,优于部分训练型方法。
原文摘要 · Abstract (English)
Accurate surgical instrument segmentation in endoscopy is crucial for computer-assisted interventions, yet remains challenging due to frequent occlusions, rapid motion, and long-term instrument re-entry. While SAM3 provides a powerful spatio-temporal framework for video object segmentation, its performance in surgical scenes is limited by indiscriminate memory updates, fixed memory capacity, and weak identity recovery after occlusions. We propose ReMeDI-SAM3, a training-free extension of SAM3, that addresses these limitations through three components: (i) relevance-aware memory filtering with a dedicated occlusion-aware memory for storing pre-occlusion frames, (ii) a piecewise interpolation scheme that expands effective memory capacity, and (iii) a feature-based re-identification module with temporal voting for reliable post-occlusion identity disambiguation. Together, these components mitigate error accumulation and enable reliable recovery after occlusions. Evaluations on EndoVis17, EndoVis18 and CholecSeg8k under a zero-shot setting show mcIoU improvements of around 5.8\%, 8\%, and 2\% respectively, over vanilla SAM3, outperforming even prior training-based approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。