arXiv:2508.05829cs.CV2025-08被引 4

针对手术视频动态复杂,提出多尺度采样与记忆分割剪枝新方法。

TSMS-SAM2: Multi-scale Temporal Sampling Augmentation and Memory-Splitting Pruning for Promptable Video Object Segmentation and Tracking in Surgical Scenarios

  • 用多时间尺度采样增强对快速运动的鲁棒性。
  • 在端口视2017和2018数据集上分别达95.24和86.73的平均骰子分数。
  • 适合需要高效精准分割的医疗视频分析场景。

可提示视频对象分割与追踪(VOST)因基础模型如SAM2的出现取得显著进展;然而,在手术视频分析中仍面临复杂运动动态和内存冗余带来的挑战。本文提出TSMS-SAM2框架,通过多时序尺度视频采样增强和记忆分割剪枝机制,提升手术视频中可提示VOST的性能。该框架在EndoVis2017和EndoVis2018数据集上分别取得95.24和86.73的最高平均骰子分数,优于已有SAM基及任务特定方法。大量消融实验验证了多尺度时间增强与记忆分割的有效性,展现了其在复杂手术场景下高效精准分割的潜力。代码将开源于https://github.com/apple1986/TSMS-SAM2。

原文摘要 · Abstract (English)

Promptable video object segmentation and tracking (VOST) has seen significant advances with the emergence of foundation models like Segment Anything Model 2 (SAM2); however, their application in surgical video analysis remains challenging due to complex motion dynamics and the redundancy of memory that impedes effective learning. In this work, we propose TSMS-SAM2, a novel framework that enhances promptable VOST in surgical videos by addressing challenges of rapid object motion and memory redundancy in SAM2. TSMS-SAM2 introduces two key strategies: multi-temporal-scale video sampling augmentation to improve robustness against motion variability, and a memory splitting and pruning mechanism that organizes and filters past frame features for more efficient and accurate segmentation. Evaluated on EndoVis2017 and EndoVis2018 datasets, TSMS-SAM2 achieved the highest mean Dice scores of 95.24 and 86.73, respectively, outperforming prior SAM-based and task-specific methods. Extensive ablation studies confirm the effectiveness of multiscale temporal augmentation and memory splitting, highlighting the framework's potential for robust, efficient segmentation in complex surgical scenarios. Our source code will be available at https://github.com/apple1986/TSMS-SAM2.

视频分割手术视频SAM2记忆剪枝

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。