arXiv:2608.20154cs.CV2026-08

用AI自动分析结直肠手术流程,跨中心通用性更强

Artificial Intelligence for Workflow Analysis in Colorectal Surgery: A Multicentric, Cross-Procedural Development and Generalization Study

论文配图:Artificial Intelligence for Workflow Analysis in Colorectal Surgery: A Multicentric, Cross-Procedural Development and Generalization Study
图 1 · 摘自论文原文
  • 用多中心数据训练视觉Transformer+时序网络,实现全自动手术阶段与步骤识别
  • 阶段识别宏F1达73.01%,步骤识别为39.82%,跨中心泛化能力优于专用模型
  • 适合做手术流程分析的AI研发者,尤其关注模型泛化与多中心部署的应用场景

微创结直肠手术(MIS-CRS)存在显著变异性和结果不一致问题。近期验证了基于视频的手术流程评估工具ColoWorkflow。然而,人工视频分析耗时长,限制了其应用。本研究提出AI-ColoWorkflow,一种用于MIS-CRS自动化手术流程分析的深度学习模型。从4个中心及一个公开数据集收集了手术视频,依据ColoWorkflow对阶段和步骤进行人工标注。采用微调的DINOv3视觉变压器提取帧级特征,并结合分层多阶段时序卷积网络,联合优化阶段与步骤识别。在合并多中心数据上训练的全局模型AI-ColoWorkflow,与中心特异性和术式特异性模型在独立测试集上进行对比。评估指标包括宏F1分数、平衡准确率、精确率和召回率。结果显示,AI-ColoWorkflow在阶段识别上达到73.01% ± 10.27的宏F1(平衡准确率73.43%),在步骤识别上为39.82% ± 7.06(平衡准确率38.65%)。除特定术式步骤识别外,全局模型在多数实验中表现优于中心或术式特异性模型。泛化分析显示,阶段识别平均F1为48.42%。AI-ColoWorkflow能可靠识别MIS-CRS阶段。单一模型在合并多中心、多术式数据上训练后,至少与中心或术式特异性模型具有同等甚至更优的泛化性能,而术式特异性步骤模型在某些术式上仍有优势,提示未来手术AI开发应考虑混合训练策略。

原文摘要 · Abstract (English)

Minimally invasive colorectal surgeries (MIS-CRS) are characterised by significant variability and inconsistent outcomes. ColoWorkflow, a tool for the video-based assessment (VBA) of MIS-CRS workflow, was recently validated. However, manual VBA is time-consuming, limiting implementation. This study presents AI-ColoWorkflow, a deep learning model for automated surgical workflow analysis across MIS-CRS. Operative videos of MIS-CRS were collected from 4 centres and a publicly available dataset. Phases and steps were manually annotated according to ColoWorkflow. A deep learning model combining a fine-tuned DINOv3 vision transformer for per-frame visual feature extraction with a hierarchical multi-stage temporal convolutional network was jointly optimized for phase and step recognition. The model trained on pooled multicentric data, namely AI-ColoWorkflow was compared against centre-specific and procedure-specific models on a held-out test set. The following metrics were used for evaluation: macro F1 score, balanced accuracy, precision, and recall. AI-ColoWorkflow achieved a macro F1 of 73.01% $\pm$ 10.27 (balanced accuracy 73.43%) for phase recognition and 39.82% $\pm$ 7.06 (balanced accuracy 38.65%) for step recognition. The global model outperformed centre- and procedure-specific models in most experiments except procedure-specific step recognition. In the generalization analysis, mean F1 was 48.42% for phase recognition. AI-ColoWorkflow can reliably recognize MIS-CRS phases. A single model trained on pooled, multicentric, multi-procedural data generalises at least as well as and often better than centre- or procedure-specific models for phase recognition in MIS-CRS, while procedure-specific step models retain advantages for certain procedure types, motivating hybrid training strategies for future surgical AI development.

手术AI流程分析多中心泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。