用AI自动分析显微手术视频,实时给出操作反馈。
An Integrated Video-AI Platform for Action-Level Microanastomosis Training and Performance Feedback

- 用Transformer模型将手术视频分割为6个动作阶段。
- 通过运动轨迹和动作统计,准确分类5项手术表现维度。
- 结合大模型实现自然语言问答,适合医学生和培训师使用。
培养显微吻合术技能需要反复练习并获得及时、针对性的反馈,但专家对长时间显微镜视频的审阅难以支持高频或分布式训练。我们提出一个集成视频-AI平台,将完整模拟手术过程转化为可检查、可交互的反馈,包含三个模块:第一,提出的Transformer模型将视频分割为六个外科操作阶段,准确率达87.66%,F1值82.86%,经流程感知优化后提升至93.62%和88.32%;第二,基于目标检测与跟踪定位器械尖端,结合运动学特征与动作统计,对五项符合NOMAT标准的性能维度进行监督分类,平均准确率76.0%,Cohen's κ值在0.63至0.93之间;第三,基于知识的大型语言模型(LLM)利用结构化输出,通过统一界面回答用户关于当前场景、操作、运动及预测表现的问题。在两站点研究中,17名参与者完成72例手术,共576次缝合操作。尽管语言接口与教学价值需前瞻性验证,结果已建立专家监督平台的技术基础,可缩短评审时间、揭示性能评估依据,并支持可扩展的形成性微外科训练。
原文摘要 · Abstract (English)
Developing microanastomosis skill requires repeated practice with timely, action-specific feedback, yet expert review of lengthy microscope videos does not scale to frequent or distributed training. We present an integrated video-AI platform that turns a complete simulated procedure into inspectable, interactive feedback through three connected modules. First, a proposed transformer segments the video into six surgical actions. Second, object detection and tracking localize instrument tips within each action; the resulting kinematic features and action statistics drive supervised classification of five NOMAT-aligned performance dimensions. Third, a grounded large language model (LLM) uses these structured outputs to answer user questions about the current scene, actions, motion, and predicted performance through a unified interface. In a two-site study, 17 participants completed 72 procedures comprising 576 suture placements. The action-segmentation module achieved 87.66\% accuracy and 82.86\% F1, increasing to 93.62\% and 88.32\% after workflow-aware refinement. The five performance classifiers achieved 76.0\% mean accuracy, with Cohen's $\kappa$ from 0.63 to 0.93. Although the language interface and educational benefit require prospective evaluation, these results establish the technical basis for an expert-supervised platform that can shorten review, expose the evidence behind performance estimates, and support scalable formative microsurgical training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。