提出可输出置信度的单目深度估计方法,提升手术视频中深度预测的可靠性。
Confidence-aware Monocular Depth Estimation for Minimally Invasive Surgery
- 用多模型集成生成像素级置信度目标,指导模型训练。
- 在自建数据集上深度估计误差降低约8%,优于基线模型。
- 推理时实时生成置信图,适合临床场景中对可靠性要求高的应用。
单目深度估计(MDE)在微创手术(MIS)中对场景理解至关重要。然而,内窥镜视频常受烟雾、反光、模糊和遮挡干扰,影响MDE精度。现有模型也不输出置信度,限制其临床可信度。本文提出一种新型置信度感知的MDE框架,包含三项贡献:(i) 校准置信度目标:通过微调的立体匹配模型集合捕捉视差方差,生成像素级置信概率;(ii) 置信度感知损失:利用像素置信度优化基线模型,使可靠像素主导训练;(iii) 推理时置信度估计:设计双卷积层头,在推理阶段预测每像素置信度,实现深度可靠性评估。在内部临床数据集(StereoKP)及公开数据集上的综合实验表明,该框架显著提升深度估计准确率,并能稳健量化预测置信度。在StereoKP数据集上,密集深度估计误差相比基线模型降低约8%。结论:该框架提升了MDE在噪声与伪影干扰下的鲁棒性,支持生成置信图,有助于提高临床应用中的可靠性。
原文摘要 · Abstract (English)
Purpose: Monocular depth estimation (MDE) is vital for scene understanding in minimally invasive surgery (MIS). However, endoscopic video sequences are often contaminated by smoke, specular reflections, blur, and occlusions, limiting the accuracy of MDE models. In addition, current MDE models do not output depth confidence, which could be a valuable tool for improving their clinical reliability. Methods: We propose a novel confidence-aware MDE framework featuring three significant contributions: (i) Calibrated confidence targets: an ensemble of fine-tuned stereo matching models is used to capture disparity variance into pixel-wise confidence probabilities; (ii) Confidence-aware loss: Baseline MDE models are optimized with confidence-aware loss functions, utilizing pixel-wise confidence probabilities such that reliable pixels dominate training; and (iii) Inference-time confidence: a confidence estimation head is proposed with two convolution layers to predict per-pixel confidence at inference, enabling assessment of depth reliability. Results: Comprehensive experimental validation across internal and public datasets demonstrates that our framework improves depth estimation accuracy and can robustly quantify the prediction's confidence. On the internal clinical endoscopic dataset (StereoKP), we improve dense depth estimation accuracy by ~8% as compared to the baseline model. Conclusion: Our confidence-aware framework enables improved accuracy of MDE models in MIS, addressing challenges posed by noise and artifacts in pre-clinical and clinical data, and allows MDE models to provide confidence maps that may be used to improve their reliability for clinical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。