arXiv:2607.16787cs.CV2026-07

用分层语义记忆提升手术视频阶段识别的准确性与连贯性

HTT-Net: Hierarchical Text-guided Transition Modeling for Surgical Video Phase Recognition

论文配图:HTT-Net: Hierarchical Text-guided Transition Modeling for Surgical Video Phase Recognition
图 1 · 摘自论文原文
  • 构建分层手术语义记忆,融合阶段内描述与阶段间过渡信息
  • 在Cholec80和LCRS-100上实现更鲁棒的阶段识别性能
  • 适合需要精准手术流程理解的临床辅助系统开发者

手术视频阶段识别是计算机辅助干预中的基础任务,有助于手术流程理解、术中引导和质量评估。尽管近期视觉时序模型取得进展,但局部视觉模糊、瞬时预测噪声以及程序语义利用不足仍导致准确且时序一致的阶段识别困难。为此,我们提出HTT-Net,一种分层文本引导的阶段转换建模网络。核心思想是将结构化手术语义知识引入阶段感知段落构建与语义优化。具体地,构建包含阶段内描述、阶段间转换描述及细粒度语义单元的分层手术语义记忆。基于该记忆,提出的过渡感知段落构建(TAS-Con)将帧级证据组织为连贯段落表示,并利用阶段间转换描述处理边界片段。此外,引入过渡感知段落校准(TAS-Calib),通过分层手术语义校准阶段感知段落表示,在无需密集帧级视觉语言融合的情况下提升视觉模糊下的判别能力。在Cholec80和LCRS-100数据集上的实验验证了HTT-Net在鲁棒手术视频阶段识别上的有效性。

原文摘要 · Abstract (English)

Surgical video phase recognition is a fundamental task in computer-assisted intervention, supporting workflow understanding, intraoperative guidance, and surgical quality assessment. Although recent visual-temporal models have achieved promising progress, accurate and temporally coherent phase recognition remains challenging due to local visual ambiguity, transient prediction noise, and insufficient use of procedural semantics. To address these challenges, we propose HTT-Net, a Hierarchical Text-guided Transition modeling Network for surgical video phase recognition. The key idea is to introduce structured surgical semantic knowledge into phase-aware segment construction and semantic refinement. Specifically, we construct a hierarchical surgical semantic memory with intra-phase descriptions, inter-phase transition descriptions, and fine-grained semantic units. Based on this memory, the proposed Transition-Aware Segment Construction (TAS-Con) organizes frame-level evidence into coherent segment representations and handles boundary clips with inter-phase transition descriptions. Furthermore, we introduce Transition-Aware Segment Calibration (TAS-Calib), which calibrates phase-aware segment representations through hierarchical surgical semantics and improves discrimination under visual ambiguity without dense frame-level vision-language fusion. Experiments on Cholec80 and LCRS-100 demonstrate the effectiveness of HTT-Net for robust surgical video phase recognition.

手术视频阶段识别语义记忆视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。