arXiv:2410.20026cs.CV2024-10被引 2

用数字孪生技术提升手术阶段识别模型的鲁棒性。

Towards Robust Algorithms for Surgical Phase Recognition via Digital Twin Representation

  • 构建数字孪生表征,分离低层视觉与高层分析
  • 在严重损坏数据上仍达80.3%准确率,优于基线3.9个百分点
  • 适合医疗影像中需高鲁棒性的智能分析场景

手术阶段识别(SPR)是手术数据分析的核心,端到端神经网络虽在基准测试中表现优异,但因训练集中的非因果关联而缺乏鲁棒性。本文提出基于数字孪生(DT)范式的框架,通过视觉基础模型实现可靠的底层场景理解,生成DT表征,并将其嵌入现有SPR模型替代原始视频输入。在Cholec80数据集上训练,评估在分布外(OOD)及受损样本上的表现。相比基线模型,本框架在高度损坏的Cholec80测试集上达到80.3%视频级准确率,在挑战性CRCD数据集上达67.9%,内部机器人手术数据集上达99.8%,分别优于基线3.9、16.8和90.9个百分点。同时发现,将DT表征作为输入增强可显著提升鲁棒性。结果表明,数字孪生表征有助于提升模型鲁棒性。未来工作将优化特征信息量并引入可解释性。

原文摘要 · Abstract (English)

Surgical phase recognition (SPR) is an integral component of surgical data science, enabling high-level surgical analysis. End-to-end trained neural networks that predict surgical phase directly from videos have shown excellent performance on benchmarks. However, these models struggle with robustness due to non-causal associations in the training set. Our goal is to improve model robustness to variations in the surgical videos by leveraging the digital twin (DT) paradigm -- an intermediary layer to separate high-level analysis (SPR) from low-level processing. As a proof of concept, we present a DT representation-based framework for SPR from videos. The framework employs vision foundation models with reliable low-level scene understanding to craft DT representation. We embed the DT representation in place of raw video inputs in the state-of-the-art SPR model. The framework is trained on the Cholec80 dataset and evaluated on out-of-distribution (OOD) and corrupted test samples. Contrary to the vulnerability of the baseline model, our framework demonstrates strong robustness on both OOD and corrupted samples, with a video-level accuracy of 80.3 on a highly corrupted Cholec80 test set, 67.9 on the challenging CRCD dataset, and 99.8 on an internal robotic surgery dataset, outperforming the baseline by 3.9, 16.8, and 90.9 respectively. We also find that using DT representation as an augmentation to the raw input can significantly improve model robustness. Our findings lend support to the thesis that DT representations are effective in enhancing model robustness. Future work will seek to improve the feature informativeness and incorporate interpretability for a more comprehensive framework.

手术识别数字孪生鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。