用元学习优化的扩散模型,提升手术阶段识别的可靠性。
Meta-SurDiff: Classification Diffusion Model Optimized by Meta Learning is Reliable for Online Surgical Phase Recognition
- 用扩散模型评估帧级识别置信度,应对视频模糊问题。
- 元学习增强分类边界鲁棒性,缓解阶段分布不均影响。
- 在5个数据集上验证,适合临床实时手术分析场景。
在线手术阶段识别因对人类健康的重要应用而受到广泛关注。尽管深度模型在捕捉手术视频长时依赖性方面取得进展,但极少关注视频中的不确定性,这对可靠识别至关重要。我们识别出两类不确定性:视频帧模糊与手术阶段分布不均,二者在手术视频中不可避免。为此,提出一种元学习优化的分类扩散模型(Meta-SurDiff),充分利用生成模型与元学习优势,实现精确的帧级分布估计。针对模糊帧导致的粗粒度识别,采用分类扩散模型在细粒度帧级评估结果置信度;针对阶段分布不均,通过元学习目标优化扩散模型,增强不同阶段分类边界的鲁棒性。在五个常用数据集(Cholec80、AutoLaparo、M2Cai16、OphNet、NurViD)上通过四种以上指标进行大量实验,验证了该方法的有效性。其中,OphNet来自眼科手术,NurViD为日常护理数据集,其余为腹腔镜手术数据集。代码将在论文录用后公开。
原文摘要 · Abstract (English)
Online surgical phase recognition has drawn great attention most recently due to its potential downstream applications closely related to human life and health. Despite deep models have made significant advances in capturing the discriminative long-term dependency of surgical videos to achieve improved recognition, they rarely account for exploring and modeling the uncertainty in surgical videos, which should be crucial for reliable online surgical phase recognition. We categorize the sources of uncertainty into two types, frame ambiguity in videos and unbalanced distribution among surgical phases, which are inevitable in surgical videos. To address this pivot issue, we introduce a meta-learning-optimized classification diffusion model (Meta-SurDiff), to take full advantage of the deep generative model and meta-learning in achieving precise frame-level distribution estimation for reliable online surgical phase recognition. For coarse recognition caused by ambiguous video frames, we employ a classification diffusion model to assess the confidence of recognition results at a finer-grained frame-level instance. For coarse recognition caused by unbalanced phase distribution, we use a meta-learning based objective to learn the diffusion model, thus enhancing the robustness of classification boundaries for different surgical phases.We establish effectiveness of Meta-SurDiff in online surgical phase recognition through extensive experiments on five widely used datasets using more than four practical metrics. The datasets include Cholec80, AutoLaparo, M2Cai16, OphNet, and NurViD, where OphNet comes from ophthalmic surgeries, NurViD is the daily care dataset, while the others come from laparoscopic surgeries. We will release the code upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。