arXiv:2509.18802cs.CV2025-09中稿 · ICRA被引 1

用光流填补手术视频标注空白,提升多任务理解精度

Surgical Video Understanding with Label Interpolation

  • 基于关键帧光流传播标签,填补非关键帧的分割缺失
  • 在5个公开数据集上平均提升12.3%的多任务指标
  • 适合需要高精度手术理解的医疗AI研究者

机器人辅助手术(RAS)已成为现代外科的关键范式,通过微创方式促进患者康复并减轻外科医生负担。为充分发挥其潜力,对术中视觉数据的精确理解至关重要。以往研究多聚焦单任务方法,但真实手术场景涉及复杂的时序动态和多样器械交互,限制了全面理解。此外,多任务学习(MTL)的有效应用需充足的像素级分割标注,而此类标注因成本高、专业性强难以获取。具体而言,长期标注如阶段和步骤覆盖每一帧,而短期标注如器械分割和动作检测仅提供关键帧,导致时空标注严重失衡。为此,本文提出一种结合光流驱动的分割标签插值与多任务学习的新框架。利用标注关键帧估计的光流,将标签传播至相邻未标注帧,从而丰富稀疏的空间监督,平衡训练中的时空信息。该集成方法提升了手术场景理解的准确性和效率,进而增强RAS的应用价值。

原文摘要 · Abstract (English)

Robot-assisted surgery (RAS) has become a critical paradigm in modern surgery, promoting patient recovery and reducing the burden on surgeons through minimally invasive approaches. To fully realize its potential, however, a precise understanding of the visual data generated during surgical procedures is essential. Previous studies have predominantly focused on single-task approaches, but real surgical scenes involve complex temporal dynamics and diverse instrument interactions that limit comprehensive understanding. Moreover, the effective application of multi-task learning (MTL) requires sufficient pixel-level segmentation data, which are difficult to obtain due to the high cost and expertise required for annotation. In particular, long-term annotations such as phases and steps are available for every frame, whereas short-term annotations such as surgical instrument segmentation and action detection are provided only for key frames, resulting in a significant temporal-spatial imbalance. To address these challenges, we propose a novel framework that combines optical flow-based segmentation label interpolation with multi-task learning. optical flow estimated from annotated key frames is used to propagate labels to adjacent unlabeled frames, thereby enriching sparse spatial supervision and balancing temporal and spatial information for training. This integration improves both the accuracy and efficiency of surgical scene understanding and, in turn, enhances the utility of RAS.

手术视频多任务学习光流插值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。