用AI自动评估外科手术技能,提升评估客观性与可扩展性。
AI-Driven Evaluation of Surgical Skill via Action Recognition
- 结合时间注意力与空间加权机制的视频变换器,精准识别手术动作。
- 在58段专家标注视频上实现93.62%的动作分割准确率,76%的评估一致性。
- 适合需标准化培训的医学生及资源有限地区的手术教学评估。
开发有效的训练与评估策略至关重要。传统外科能力评估依赖专家现场观察或事后录像分析,但此类方法主观性强、评分者间差异大,且耗时耗力,难以在低收入和中等收入国家推广。为此,本文提出一种基于AI的微血管吻合术性能自动化评估框架。系统采用改进的TimeSformer视频变换器架构,引入分层时间注意力与加权空间注意力机制,实现手术视频中动作的精准识别;并结合基于YOLO的对象检测与追踪方法提取细粒度运动特征,分析器械运动轨迹。评估涵盖五个方面:整体动作执行、关键步骤运动质量及器械通用操作。在包含58段专家标注视频的数据集上验证,系统在帧级动作分割上达到87.7%准确率,经后处理提升至93.62%;在所有技能维度上平均分类准确率达76%,复现专家评估结果。该系统具备客观、一致、可解释的反馈能力,为外科教育提供标准化、数据驱动的训练与评估支持。
原文摘要 · Abstract (English)
The development of effective training and evaluation strategies is critical. Conventional methods for assessing surgical proficiency typically rely on expert supervision, either through onsite observation or retrospective analysis of recorded procedures. However, these approaches are inherently subjective, susceptible to inter-rater variability, and require substantial time and effort from expert surgeons. These demands are often impractical in low- and middle-income countries, thereby limiting the scalability and consistency of such methods across training programs. To address these limitations, we propose a novel AI-driven framework for the automated assessment of microanastomosis performance. The system integrates a video transformer architecture based on TimeSformer, improved with hierarchical temporal attention and weighted spatial attention mechanisms, to achieve accurate action recognition within surgical videos. Fine-grained motion features are then extracted using a YOLO-based object detection and tracking method, allowing for detailed analysis of instrument kinematics. Performance is evaluated along five aspects of microanastomosis skill, including overall action execution, motion quality during procedure-critical actions, and general instrument handling. Experimental validation using a dataset of 58 expert-annotated videos demonstrates the effectiveness of the system, achieving 87.7% frame-level accuracy in action segmentation that increased to 93.62% with post-processing, and an average classification accuracy of 76% in replicating expert assessments across all skill aspects. These findings highlight the system's potential to provide objective, consistent, and interpretable feedback, thereby enabling more standardized, data-driven training and evaluation in surgical education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。