用时间可切换的师生学习法,解决微创手术视频立体匹配的伪标签不稳问题。
TiS-TSL: Image-Label Supervised Surgical Video Stereo Matching via Time-Switchable Teacher-Student Learning
- 设计三模式统一模型,实现前后向视频预测与图像预测融合。
- 通过双向时序一致性检测,减少伪标签噪声并消除帧间闪烁。
- 适合需要高精度立体视觉的手术导航与增强现实系统开发者。
微创手术中的立体匹配对下一代导航与增强现实至关重要,但因解剖限制,密集视差标注几乎不可行,通常仅能获取少数进入体腔前的图像级标签。教师-学生学习(TSL)通过教师模型从大量未标注手术视频中生成伪标签与置信图,提供可行方案。然而,现有TSL方法局限于图像级监督,仅有空间置信度,缺乏时序一致性估计,导致视差预测不稳定且出现严重帧间闪烁。为此,本文提出TiS-TSL——一种面向极少量监督下的视频立体匹配的时间可切换师生学习框架。核心是统一模型,支持图像预测(IP)、前向视频预测(FVP)和后向视频预测(BVP)三种模式,实现灵活的时序建模。采用两阶段训练策略:首先在图像到视频(I2V)阶段将稀疏图像级知识迁移至时序建模;再在视频到视频(V2V)阶段通过比较前向与后向预测,计算双向时空一致性,识别不可靠区域,过滤噪声视频级伪标签,并强制时序一致性。在两个公开数据集上的实验表明,相比其他基于图像的先进方法,TiS-TSL在TEPE和EPE上分别提升至少2.11%和4.54%。
原文摘要 · Abstract (English)
Stereo matching in minimally invasive surgery (MIS) is essential for next-generation navigation and augmented reality. Yet, dense disparity supervision is nearly impossible due to anatomical constraints, typically limiting annotations to only a few image-level labels acquired before the endoscope enters deep body cavities. Teacher-Student Learning (TSL) offers a promising solution by leveraging a teacher trained on sparse labels to generate pseudo labels and associated confidence maps from abundant unlabeled surgical videos. However, existing TSL methods are confined to image-level supervision, providing only spatial confidence and lacking temporal consistency estimation. This absence of spatio-temporal reliability results in unstable disparity predictions and severe flickering artifacts across video frames. To overcome these challenges, we propose TiS-TSL, a novel time-switchable teacher-student learning framework for video stereo matching under minimal supervision. At its core is a unified model that operates in three distinct modes: Image-Prediction (IP), Forward Video-Prediction (FVP), and Backward Video-Prediction (BVP), enabling flexible temporal modeling within a single architecture. Enabled by this unified model, TiS-TSL adopts a two-stage learning strategy. The Image-to-Video (I2V) stage transfers sparse image-level knowledge to initialize temporal modeling. The subsequent Video-to-Video (V2V) stage refines temporal disparity predictions by comparing forward and backward predictions to calculate bidirectional spatio-temporal consistency. This consistency identifies unreliable regions across frames, filters noisy video-level pseudo labels, and enforces temporal coherence. Experimental results on two public datasets demonstrate that TiS-TSL exceeds other image-based state-of-the-arts by improving TEPE and EPE by at least 2.11% and 4.54%, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。