利用立体与时间上下文提升手术器械分割精度
LACOSTE: Exploiting stereo and temporal contexts for surgical instrument segmentation
- 通过立体和时间上下文建模增强深度感知特征
- 在三个数据集上达到或超越现有最佳效果
- 适合关注微创手术视觉分析的研究者
手术器械分割对微创手术及其应用至关重要。以往方法多基于单帧实例分割,忽视了手术视频固有的时空属性,导致在运动和视角变化下鲁棒性不足。本文提出LACOSTE模型,利用立体与时间上下文中的位置无关上下文信息,提升分割性能。以查询式分割模型为核心,设计三个增强模块:首先,设计视差引导的特征传播模块,显式增强深度感知特征;为适应仅单目视频场景,引入伪立体方案生成互补右图像;其次,提出立体-时间集合分类器,统一聚合时空上下文以实现综合预测,缓解瞬时失败;最后,设计位置无关分类器,解耦掩码预测中的位置偏差,强化特征语义。在三个公开数据集上验证,包括两个来自EndoVis挑战赛的基准数据集和真实前列腺根治术数据集GraSP。实验表明,本方法在各项指标上均表现优异,持续达到或优于现有最先进水平。
原文摘要 · Abstract (English)
Surgical instrument segmentation is instrumental to minimally invasive surgeries and related applications. Most previous methods formulate this task as single-frame-based instance segmentation while ignoring the natural temporal and stereo attributes of a surgical video. As a result, these methods are less robust against the appearance variation through temporal motion and view change. In this work, we propose a novel LACOSTE model that exploits Location-Agnostic COntexts in Stereo and TEmporal images for improved surgical instrument segmentation. Leveraging a query-based segmentation model as core, we design three performance-enhancing modules. Firstly, we design a disparity-guided feature propagation module to enhance depth-aware features explicitly. To generalize well for even only a monocular video, we apply a pseudo stereo scheme to generate complementary right images. Secondly, we propose a stereo-temporal set classifier, which aggregates stereo-temporal contexts in a universal way for making a consolidated prediction and mitigates transient failures. Finally, we propose a location-agnostic classifier to decouple the location bias from mask prediction and enhance the feature semantics. We extensively validate our approach on three public surgical video datasets, including two benchmarks from EndoVis Challenges and one real radical prostatectomy surgery dataset GraSP. Experimental results demonstrate the promising performances of our method, which consistently achieves comparable or favorable results with previous state-of-the-art approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。