对比医生与AI在复杂手术阶段识别中的表现,发现两者在有时间信息时能力相当。
Surgeons vs. Computer Vision: A comparative analysis on surgical phase recognition capabilities
- 用视频片段和视觉标记提升阶段识别准确率
- 专家比新手更擅长区分手术阶段,且信心更高
- 时间上下文让医生和AI表现都明显提升
自动化手术阶段识别(SPR)利用人工智能将手术流程分割为关键事件,是高效视频回放、手术教学与技能评估的基础。以往研究多集中于短而线性的手术,未探讨时间上下文对专家判断的影响。本研究以高度非线性的机器人辅助部分肾切除术(RAPN)为例,招募不同经验的泌尿科医生,在自研网页平台上基于单帧图像和视频片段标注手术阶段,并报告其置信度及所依赖的视觉线索。随后在Cholec80数据集上训练并对比无/有时间上下文的AI模型,再在RAPN数据集上进行微调。结果显示:视频片段和特定视觉地标均显著提升各组识别准确率;医生普遍信心高,专家优于新手;AI模型性能与医生相当,加入时间上下文后进一步提升。结论:无论人类还是计算机视觉,阶段识别均为复杂任务,但当提供相同上下文时表现相近,时间信息对提升性能至关重要。手术器械与器官是人类判断的关键视觉地标,也应成为未来自动识别的核心依据。
原文摘要 · Abstract (English)
Purpose: Automated Surgical Phase Recognition (SPR) uses Artificial Intelligence (AI) to segment the surgical workflow into its key events, functioning as a building block for efficient video review, surgical education as well as skill assessment. Previous research has focused on short and linear surgical procedures and has not explored if temporal context influences experts' ability to better classify surgical phases. This research addresses these gaps, focusing on Robot-Assisted Partial Nephrectomy (RAPN) as a highly non-linear procedure. Methods: Urologists of varying expertise were grouped and tasked to indicate the surgical phase for RAPN on both single frames and video snippets using a custom-made web platform. Participants reported their confidence levels and the visual landmarks used in their decision-making. AI architectures without and with temporal context as trained and benchmarked on the Cholec80 dataset were subsequently trained on this RAPN dataset. Results: Video snippets and presence of specific visual landmarks improved phase classification accuracy across all groups. Surgeons displayed high confidence in their classifications and outperformed novices, who struggled discriminating phases. The performance of the AI models is comparable to the surgeons in the survey, with improvements when temporal context was incorporated in both cases. Conclusion: SPR is an inherently complex task for expert surgeons and computer vision, where both perform equally well when given the same context. Performance increases when temporal information is provided. Surgical tools and organs form the key landmarks for human interpretation and are expected to shape the future of automated SPR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。