arXiv:2606.26634cs.CV2026-06

用光流与零样本分割,从稀疏标注生成稠密伪标签,提升手术多任务学习效果

Temporally Consistent Label Interpolation for Robust Surgical Multi-Task Learning under Challenging Conditions

论文配图:Temporally Consistent Label Interpolation for Robust Surgical Multi-Task Learning under Challenging Conditions
图 1 · 摘自论文原文
  • 通过光流+零样本分割,从关键帧稀疏标注生成连续帧伪标签
  • 在3个手术数据集上显著提升相位、步骤、器械分割等任务性能
  • 适合需要稠密时间监督但标注成本高的医疗视觉研究者

手术场景理解的多任务学习受标注粒度不匹配制约:时间任务(如阶段识别、步骤识别)需逐帧精细标注,而空间任务(如器械分割、动作识别)因标注成本高,仅在关键帧上有稀疏标注。这种监督失衡阻碍共享表征学习。为此,我们提出FAROS框架,结合零样本分割与光流估计,克服遮挡、烟雾、运动模糊等挑战下的外观传播失效问题,从稀疏关键帧标注生成时间一致的稠密伪标签。将生成的器械掩码与动作标签融入统一的Transformer多任务框架,联合优化相位识别、步骤识别、预测、器械分割和动作识别,实现密集时间监督与稀疏空间监督的平衡优化。FAROS在DAVIS 2017稀疏真值协议下验证了跨域鲁棒性;在GraSP、MISAW和AutoLaparo基准上进一步证明其显著提升跨任务表征学习能力,增强整体手术场景理解性能。

原文摘要 · Abstract (English)

Effective multi-task learning for surgical scene understanding is fundamentally hindered by annotation granularity mismatch; temporal workflow tasks such as phase recognition, step recognition and anticipation benefit from dense frame-level supervision, whereas pixel-level spatial tasks including instrument segmentation and action recognition are only sparsely annotated on selected keyframes due to prohibitive labeling costs. This supervision imbalance undermines shared representation learning and limits joint optimization across heterogeneous surgical tasks. To address this, we propose Flow-guided Annotation for Robust Operating Scenes (FAROS), a flow-guided label interpolation framework, that combines zero-shot segmentation-based mask propagation with optical flow estimation to overcome the limitations of appearance-based propagation under challenging surgical conditions such as occlusion, smoke, and motion blur, generating temporally consistent dense pseudo labels from sparse keyframe annotations. The densified instrument masks and action labels are integrated into a unified Transformer-based multi-task framework that jointly learns surgical phase recognition, step recognition, anticipation, instrument segmentation, and action recognition, enabling balanced optimization between dense temporal supervision and sparse spatial supervision. The label interpolation quality of FAROS is first validated on the DAVIS 2017 benchmark under a sparse ground-truth protocol, confirming robust propagation beyond the surgical domain. Extensive experiments on GraSP, MISAW, and AutoLaparo benchmarks further demonstrate that FAROS significantly improves cross-task representation learning and enhances holistic surgical scene understanding performance across spatio-temporal tasks.

多任务学习手术视觉伪标签光流

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。