arXiv:2502.20934cs.CVcs.AI2025-02中稿 · publication in the…被引 2

低帧率采样会误导模型评估,真实手术场景需高帧率测试

Revisiting the Evaluation Bias Introduced by Frame Sampling Strategies in Surgical Video Segmentation Using SAM2

  • 用SAM2分析胆囊切除术视频,对比不同帧率下分割表现
  • 低帧率(1 FPS)因平滑效应虚假提升性能,实际稳定性差
  • 医生护士偏好高帧率分割,真实手术应实时评估每帧

实时视频分割为智能手术提供术中引导,可识别器械与解剖结构。尽管研究兴趣上升,不同数据集的标注协议差异大:有的提供逐帧密集标注,有的仅在低帧率(如1 FPS)下稀疏采样。本研究以胆囊切除术为例,考察标注密度与帧率对零样本分割模型(使用SAM2)评估的影响。令人意外的是,在传统稀疏评估设置下,低帧率反而看似优于高帧率,因其平滑效应掩盖了时间不一致性。但在真实流式条件下,高帧率能显著提升动态物体(如手术抓钳)的分割稳定性。我们对外科医生、护士及机器学习工程师进行调研,结果一致显示更偏好高帧率分割叠加效果。研究揭示了评估偏差风险,强调手术视频人工智能需采用时序公平的基准测试。

原文摘要 · Abstract (English)

Real-time video segmentation is a promising opportunity for AI-assisted surgery, offering intraoperative guidance by identifying tools and anatomical structures. Despite growing interest in surgical video segmentation, annotation protocols vary widely across datasets -- some provide dense, frame-by-frame labels, while others rely on sparse annotations sampled at low frame rates such as 1 FPS. In this study, we investigate how such inconsistencies in annotation density and frame rate sampling influence the evaluation of zero-shot segmentation models, using SAM2 as a case study for cholecystectomy procedures. Surprisingly, we find that under conventional sparse evaluation settings, lower frame rates can appear to outperform higher ones due to a smoothing effect that conceals temporal inconsistencies. However, when assessed under real-time streaming conditions, higher frame rates yield superior segmentation stability, particularly for dynamic objects like surgical graspers. To understand how these differences align with human perception, we conducted a survey among surgeons, nurses, and machine learning engineers and found that participants consistently preferred high-FPS segmentation overlays, reinforcing the importance of evaluating every frame in real-time applications rather than relying on sparse sampling strategies. Our findings highlight the risk of evaluation bias that is introduced by inconsistent dataset protocols and bring attention to the need for temporally fair benchmarking in surgical video AI.

视频分割手术AI评估偏差SAM2

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。