arXiv:2601.12312cs.CV2026-01

提出统一时空框架,提升手术流程细粒度理解能力

CurConMix+: A Unified Spatio-Temporal Framework for Hierarchical Surgical Workflow Understanding

  • 采用课程引导对比学习与特征混合,增强空间特征判别性
  • 多分辨率时序变换器融合多尺度时序信息,动态平衡时空线索
  • 新构建标注精细的LLS48数据集,支持跨层级泛化研究

手术行为三元组识别旨在通过建模器械、操作和解剖目标间的交互,理解精细的手术行为。尽管对流程分析与技能评估具有重要临床价值,但严重类别不平衡、细微视觉差异及三元组成分间语义依赖性阻碍了进展。现有方法通常仅解决部分挑战,难以实现整体理解。本文在CurConMix空间表示框架基础上,提出CurConMix+,引入课程引导对比学习策略,结合结构化难样本采样与特征级混合,逐步学习更具区分性的相关特征。其时序扩展模块——多分辨率时序变换器(MRTT),通过自适应融合多尺度时序特征,实现鲁棒且上下文感知的理解。此外,本文构建了新基准LLS48,针对复杂腹腔镜左外侧切除术提供步骤、任务、操作级分层标注。在CholecT45和LLS48上的实验表明,CurConMix+不仅显著优于现有先进方法,在三元组识别上表现更优,且具备强跨层级泛化能力,其细粒度特征可有效迁移到更高层级的阶段与步骤识别任务中。该框架与数据集共同为层次感知、可复现、可解释的手术流程理解提供统一基础。代码与数据集将公开于GitHub以促进研究。

原文摘要 · Abstract (English)

Surgical action triplet recognition aims to understand fine-grained surgical behaviors by modeling the interactions among instruments, actions, and anatomical targets. Despite its clinical importance for workflow analysis and skill assessment, progress has been hindered by severe class imbalance, subtle visual variations, and the semantic interdependence among triplet components. Existing approaches often address only a subset of these challenges rather than tackling them jointly, which limits their ability to form a holistic understanding. This study builds upon CurConMix, a spatial representation framework. At its core, a curriculum-guided contrastive learning strategy learns discriminative and progressively correlated features, further enhanced by structured hard-pair sampling and feature-level mixup. Its temporal extension, CurConMix+, integrates a Multi-Resolution Temporal Transformer (MRTT) that achieves robust, context-aware understanding by adaptively fusing multi-scale temporal features and dynamically balancing spatio-temporal cues. Furthermore, we introduce LLS48, a new, hierarchically annotated benchmark for complex laparoscopic left lateral sectionectomy, providing step-, task-, and action-level annotations. Extensive experiments on CholecT45 and LLS48 demonstrate that CurConMix+ not only outperforms state-of-the-art approaches in triplet recognition, but also exhibits strong cross-level generalization, as its fine-grained features effectively transfer to higher-level phase and step recognition tasks. Together, the framework and dataset provide a unified foundation for hierarchy-aware, reproducible, and interpretable surgical workflow understanding. The code and dataset will be publicly released on GitHub to facilitate reproducibility and further research.

手术理解三元组识别多模态视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。