arXiv:2502.15184cs.CV2025-02中稿 · the IEEE TCSVT被引 10

提出分层上下文变换器,实现手术场景多级语义理解

Hierarchical Context Transformer for Multi-level Semantic Scene Understanding

  • 设计分层关系聚合模块,统一建模多层级任务间关系
  • 在白内障和内窥镜数据集上显著超越现有方法
  • 适合关注手术理解与医疗AI的开发者与研究者

全面且明确的手术场景理解对开发术中智能辅助系统至关重要。然而,目前缺乏系统性分析来实现分层手术场景理解。本文将阶段识别、步骤识别、动作与器械检测任务统一为多级语义场景理解(MSSU),提出新型分层上下文变换器(HCT)网络,并深入探索不同层级任务间的关联。设计分层关系聚合模块(HRAM),同步融合多层级交互信息以增强任务特定特征。为进一步提升表示学习能力,引入跨任务对比学习(ICL),通过吸收其他任务的互补信息指导模型学习任务特异性特征。针对Transformer计算开销问题,提出HCT+,集成空间与时间适配器,在大幅减少可调参数的前提下保持优异性能。在自建白内障数据集及公开的内窥镜PSI-AVA数据集上的大量实验表明,该方法性能显著优于现有最先进方法。代码已开源。

原文摘要 · Abstract (English)

A comprehensive and explicit understanding of surgical scenes plays a vital role in developing context-aware computer-assisted systems in the operating theatre. However, few works provide systematical analysis to enable hierarchical surgical scene understanding. In this work, we propose to represent the tasks set [phase recognition --> step recognition --> action and instrument detection] as multi-level semantic scene understanding (MSSU). For this target, we propose a novel hierarchical context transformer (HCT) network and thoroughly explore the relations across the different level tasks. Specifically, a hierarchical relation aggregation module (HRAM) is designed to concurrently relate entries inside multi-level interaction information and then augment task-specific features. To further boost the representation learning of the different tasks, inter-task contrastive learning (ICL) is presented to guide the model to learn task-wise features via absorbing complementary information from other tasks. Furthermore, considering the computational costs of the transformer, we propose HCT+ to integrate the spatial and temporal adapter to access competitive performance on substantially fewer tunable parameters. Extensive experiments on our cataract dataset and a publicly available endoscopic PSI-AVA dataset demonstrate the outstanding performance of our method, consistently exceeding the state-of-the-art methods by a large margin. The code is available at https://github.com/Aurora-hao/HCT.

手术理解分层建模视觉Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。