arXiv:2504.18756cs.CV2025-04被引 2

精准分割手术视频中模糊边界的动作,提升训练与评估效果。

Multi-Stage Boundary-Aware Transformer Network for Action Segmentation in Untrimmed Surgical Videos

  • 分阶段设计注意力机制,捕捉动作边界与长短时序变化。
  • 在三个数据集上达25%和50%阈值下最优F1分数。
  • 创新边界加权策略,结合上下文信息精确定位动作起止点。

理解手术流程中的动作对评估术后结果、提升手术培训与效率至关重要。由于不同外科医生因经验与偏好导致操作差异,长序列动作的识别与分割面临边界模糊、起止点不清晰的挑战。传统模型如MS-TCN依赖大感受野,易引发过分割或欠分割问题。为此,我们提出多阶段边界感知变压器网络(MSBATN),采用层次化滑动窗口注意力机制,有效应对动作持续时间差异与细微过渡。MSBATN引入统一损失函数,将动作分类与边界检测联合优化。不同于传统二值边界检测,其创新的边界加权机制利用上下文信息实现精准定位。在三个具有挑战性的手术数据集上的实验表明,该方法在25%和50%阈值下取得最佳F1得分,其他指标表现也具竞争力。

原文摘要 · Abstract (English)

Understanding actions within surgical workflows is critical for evaluating post-operative outcomes and enhancing surgical training and efficiency. Capturing and analyzing long sequences of actions in surgical settings is challenging due to the inherent variability in individual surgeon approaches, which are shaped by their expertise and preferences. This variability complicates the identification and segmentation of distinct actions with ambiguous boundary start and end points. The traditional models, such as MS-TCN, which rely on large receptive fields, that causes over-segmentation, or under-segmentation, where distinct actions are incorrectly aligned. To address these challenges, we propose the Multi-Stage Boundary-Aware Transformer Network (MSBATN) with hierarchical sliding window attention to improve action segmentation. Our approach effectively manages the complexity of varying action durations and subtle transitions by accurately identifying start and end action boundaries in untrimmed surgical videos. MSBATN introduces a novel unified loss function that optimises action classification and boundary detection as interconnected tasks. Unlike conventional binary boundary detection methods, our innovative boundary weighing mechanism leverages contextual information to precisely identify action boundaries. Extensive experiments on three challenging surgical datasets demonstrate that MSBATN achieves state-of-the-art performance, with superior F1 scores at 25% and 50%. thresholds and competitive results across other metrics.

动作分割手术视频边界感知Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。