用视频自动识别手术动作并预测恢复效果,准确率媲美医生。
End to End AI System for Surgical Gesture Sequence Recognition and Clinical Outcome Prediction
- 用变换器模型分析手术视频,精准识别每段约2秒的精细操作
- 提取动作频率、时长等特征,预测术后恢复效果准确率达0.79
- 能发现影响性功能恢复的关键操作模式,适合临床反馈与决策支持
术中行为的细粒度分析及其对患者预后的影响仍是长期挑战。我们提出端到端的帧到结果(F2O)系统,将组织分离视频转化为动作序列,并揭示其与术后结局的关联。该系统基于变换器进行时空建模与逐帧分类,在机器人辅助根治性前列腺切除术的神经保留步骤中,实现了帧级AUC 0.80、视频级AUC 0.81。F2O提取的动作频率、持续时间及转换特征,预测术后结局的准确率为0.79,与人工标注(0.75)相当(95%置信区间重叠)。在25个共享特征中,效应方向一致,差异小(~0.07),相关性极强(r = 0.96,p < 1e-14)。F2O还识别出与勃起功能恢复相关的关键模式,如组织剥离时间延长、能量使用减少。该系统可实现自动可解释评估,为数据驱动的手术反馈与前瞻性临床决策支持奠定基础。
原文摘要 · Abstract (English)
Fine-grained analysis of intraoperative behavior and its impact on patient outcomes remain a longstanding challenge. We present Frame-to-Outcome (F2O), an end-to-end system that translates tissue dissection videos into gesture sequences and uncovers patterns associated with postoperative outcomes. Leveraging transformer-based spatial and temporal modeling and frame-wise classification, F2O robustly detects consecutive short (~2 seconds) gestures in the nerve-sparing step of robot-assisted radical prostatectomy (AUC: 0.80 frame-level; 0.81 video-level). F2O-derived features (gesture frequency, duration, and transitions) predicted postoperative outcomes with accuracy comparable to human annotations (0.79 vs. 0.75; overlapping 95% CI). Across 25 shared features, effect size directions were concordant with small differences (~ 0.07), and strong correlation (r = 0.96, p < 1e-14). F2O also captured key patterns linked to erectile function recovery, including prolonged tissue peeling and reduced energy use. By enabling automatic interpretable assessment, F2O establishes a foundation for data-driven surgical feedback and prospective clinical decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。