arXiv:2605.23653cs.CV2026-05

用3D手部重建自动评估开刀技能,还能解释哪里好哪里差。

ExpOS: Explainable Open-Surgery Skills Assessment Using 3D Hand Reconstruction

论文配图:ExpOS: Explainable Open-Surgery Skills Assessment Using 3D Hand Reconstruction
图 1 · 摘自论文原文
  • 从动作数据中学习关键时间片段和行为模式,无需专家定义指标。
  • 在221段学生手术视频上,对筋膜缝合的评估相关性达0.778。
  • 通过注意力机制和特征分析,提供可解释的反馈,适合训练使用。

及时透明的反馈对有效外科训练至关重要,但当前评估仍依赖专家观察,难以扩展且限制自主练习。我们提出ExpOS,一种基于数据驱动的可解释开刀技能评估框架,实现自动、反馈导向的评价。该方法不依赖专家设定指标,而是直接从运动数据中学习具有区分性的时序模式,并识别出最能预测技能水平的时间段与行为。我们在221段医学生完成三项开刀任务的视频上训练和评估该方法。每帧提取手部姿态与器械检测结果,生成运动学描述符与全局运动统计量。采用带注意力池化的时序卷积主干模型,建模时空手-器械动态,生成帧级重要性图。这些表示与全局运动统计量融合,用于预测技能水平并提供可解释反馈。ExpOS通过注意力权重识别关键事件发生时刻,通过全局特征分析揭示影响预测的主要运动特征,实现多层次可解释性。在各项任务中,该框架与专家评分呈现强相关性,最佳表现于筋膜缝合任务(r = 0.778,R² = 0.74)。结果表明,结合弱监督时序重要性学习与可解释运动统计量,可实现可扩展且可操作的外科技能评估。

原文摘要 · Abstract (English)

Timely and transparent feedback is essential for effective surgical training, yet current assessment remains dependent on expert observation, limiting scalability and opportunities for autonomous practice. We present ExpOS, an explainable framework for data-driven assessment of open-surgery skills designed to enable automatic, feedback-oriented evaluation. Rather than relying on expert-defined metrics, ExpOS learns discriminative temporal patterns directly from motion data and identifies the segments and behaviors most predictive of skill level. We trained and evaluated the method on 221 videos of medical students performing three open-surgery tasks. Hand poses and tool detections were extracted from each frame to derive kinematic descriptors and global motion statistics. Spatiotemporal hand-tool dynamics were modeled using a temporal convolutional backbone with attention-based pooling to generate frame-level importance maps. These representations were fused with global motion statistics to predict skill level and to provide interpretable feedback. ExpOS provides multi-level explainability by identifying when informative events occur through attention weights and which motion characteristics most influence predictions through global feature analysis. Across tasks, the framework achieved strong correlation with expert ratings, with best performance on fascial closure (r = 0.778, R2 = 0.74). These results demonstrate that combining weakly-supervised temporal importance learning with interpretable motion statistics enables scalable and actionable surgical skill assessment.

手术评估可解释性3D重建动作识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。