用工具运动分析精简手术视频,仅用0.58%帧数实现91%准确率
Extended KAFR: A kinematic-adaptive paradigm for the efficient analysis of surgical video

- 基于工具位移和速度变化动态筛选关键帧
- 仅用0.58%帧数达到91.0%的相位分类F1分数
- 适合需要高效处理长时手术视频的研究者
人工智能在手术视频分析中用于阶段分割、技能评估与流程优化。手术录像常长达一至数小时,带来巨大计算负担。我们此前提出的运动自适应帧识别(KAFR)方法在机器人手术中有效识别关键帧并过滤冗余内容。本研究评估KAFR在腹腔镜手术中的泛化能力,使用包含80例胆囊切除术的Cholec80数据集,该数据集标注了七个手术阶段。KAFR分三步:首先用微调的YOLO模型检测并分割手术工具;其次根据工具位移或速度变化自适应选择帧;最后用X3D模型对选中帧进行阶段分类。结果显示,仅使用0.58%的帧数,KAFR即取得91.0%的F1分数,相较常规4%采样减少约七倍,性能接近LoViT(90.2%)和Trans-SVNet(89.7%)。这证明基于运动特征的帧选择可有效迁移至挑战性强的腹腔镜环境。
原文摘要 · Abstract (English)
Artificial Intelligence is increasingly applied to surgical video analysis for phase segmentation, skill assessment, and workflow optimization. A key challenge is the length of surgical recordings, often one to several hours, creating substantial computational burden. We previously developed Kinematics-Adaptive Frame Recognition (KAFR) for robotic surgery, showing that tracking tool motion effectively identifies informative frames while filtering redundant content. However, laparoscopic surgery introduces additional challenges: manual camera control causes frequent motion artifacts, and image quality is generally lower than robotic systems. This study evaluates whether KAFR generalizes to laparoscopic surgery using the Cholec80 benchmark, comprising 80 laparoscopic cholecystectomy procedures annotated for seven surgical phases. KAFR operates in three stages: a fine-tuned YOLO model detects and segments surgical tools; frames are adaptively selected based on tool displacement or velocity variation; and an X3D model classifies selected frames into surgical phases. KAFR achieved a 91.0\% F1 score using only 0.58\% of frames for phase classification, representing an approximately seven-fold reduction compared to typical 4\% frame sampling, while maintaining performance comparable to LoViT (90.2\%) and Trans-SVNet (89.7\%). These results demonstrate that kinematics-based frame selection transfers effectively to the challenging laparoscopic environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。