用可解释的动态图原型建模手术流程,少样本下仍精准可靠
ProtoFlow: Interpretable and Robust Surgical Workflow Modeling with Learned Dynamic Scene Graph Prototypes
- 通过自监督预训练+原型精调,学习临床有意义的手术模式
- 在仅1段视频训练时仍保持高准确率,少样本性能超越基线
- 原型可解析识别手术技巧与异常,适合手术教学与实时辅助
目的:精细的手术识别对推进人工智能辅助手术至关重要,但受限于高标注成本、数据稀缺及缺乏可解释模型。尽管场景图能结构化抽象手术事件,其潜力尚未充分挖掘。本文提出ProtoFlow,一种基于学习的动态场景图原型框架,可解释且鲁棒地建模复杂手术流程。方法:ProtoFlow采用图神经网络(GNN)编码器-解码器架构,结合自监督预训练以实现丰富表征学习,并引入基于原型的微调阶段,发现并优化核心原型,捕捉重复出现且具有临床意义的手术交互模式,为流程分析提供可解释基础。结果:在细粒度CAT-SG数据集上评估,ProtoFlow不仅整体准确率优于标准GNN基线,且在低数据、少样本场景中表现卓越,仅需1段手术视频训练即维持强性能。定性分析表明,学习到的原型成功识别不同手术子技术,并清晰揭示流程偏差与罕见并发症。结论:通过融合鲁棒表征学习与内在可解释性,ProtoFlow推动了更透明、可靠、数据高效的人工智能系统发展,加速其在手术培训、实时决策支持与流程优化中的临床应用。
原文摘要 · Abstract (English)
Purpose: Detailed surgical recognition is critical for advancing AI-assisted surgery, yet progress is hampered by high annotation costs, data scarcity, and a lack of interpretable models. While scene graphs offer a structured abstraction of surgical events, their full potential remains untapped. In this work, we introduce ProtoFlow, a novel framework that learns dynamic scene graph prototypes to model complex surgical workflows in an interpretable and robust manner. Methods: ProtoFlow leverages a graph neural network (GNN) encoder-decoder architecture that combines self-supervised pretraining for rich representation learning with a prototype-based fine-tuning stage. This process discovers and refines core prototypes that encapsulate recurring, clinically meaningful patterns of surgical interaction, forming an explainable foundation for workflow analysis. Results: We evaluate our approach on the fine-grained CAT-SG dataset. ProtoFlow not only outperforms standard GNN baselines in overall accuracy but also demonstrates exceptional robustness in limited-data, few-shot scenarios, maintaining strong performance when trained on as few as one surgical video. Our qualitative analyses further show that the learned prototypes successfully identify distinct surgical sub-techniques and provide clear, interpretable insights into workflow deviations and rare complications. Conclusion: By uniting robust representation learning with inherent explainability, ProtoFlow represents a significant step toward developing more transparent, reliable, and data-efficient AI systems, accelerating their potential for clinical adoption in surgical training, real-time decision support, and workflow optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。