分析手术视频中点追踪的失败模式,提升精准度
When Tracking Fails: Analyzing Failure Modes of SAM2 for Point-Based Tracking in Surgical Videos
- 对比点输入与掩码初始化,评估跟踪效果差异
- 工具类目标点追踪表现良好,解剖结构易失败
- 提供点选择策略,适合临床视频分析场景
视频对象分割模型如SAM2在仅需少量用户输入的情况下,可实现手术视频的零样本跟踪。点基追踪虽效率高、成本低,但在复杂手术环境中其可靠性与失败原因尚不明确。本文系统分析了腹腔镜胆囊切除术视频中三种目标(胆囊、抓钳、L型电刀)的点基追踪失败模式。结果表明:点追踪对器械类目标表现良好,但对解剖结构目标始终表现较差,主要因组织相似性和边界模糊导致。通过定性分析,揭示影响追踪的关键因素,并提出若干可操作的点选择与放置建议,以优化手术视频分析中的追踪性能。
原文摘要 · Abstract (English)
Video object segmentation (VOS) models such as SAM2 offer promising zero-shot tracking capabilities for surgical videos using minimal user input. Among the available input types, point-based tracking offers an efficient and low-cost alternative, yet its reliability and failure cases in complex surgical environments are not well understood. In this work, we systematically analyze the failure modes of point-based tracking in laparoscopic cholecystectomy videos. Focusing on three surgical targets, the gallbladder, grasper, and L-hook electrocautery, we compare the performance of point-based tracking with segmentation mask initialization. Our results show that point-based tracking is competitive for surgical tools but consistently underperforms for anatomical targets, where tissue similarity and ambiguous boundaries lead to failure. Through qualitative analysis, we reveal key factors influencing tracking outcomes and provide several actionable recommendations for selecting and placing tracking points to improve performance in surgical video analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。