arXiv:2606.31245cs.CV2026-06

用双曲空间建模手术视频语言的层级结构,提升跨机构识别效果

HyperVLP: Enhancing Hierarchical Surgical Video-Language Pre-training in Hyperbolic Space

论文配图:HyperVLP: Enhancing Hierarchical Surgical Video-Language Pre-training in Hyperbolic Space
图 1 · 摘自论文原文
  • 在双曲空间中显式保持手术步骤的层级关系
  • 在多个数据集上实现零样本和少样本阶段识别显著提升
  • 适合需要跨机构泛化的手术视觉理解研究者

手术视觉-语言基础模型通常使用教学材料(如手术讲座视频)将语言编码的手术知识迁移到视觉表征中。这些知识具有多维且分层的特性:细粒度操作线索出现在叙述中,中等层次的关键步骤由小节标题总结,全局流程上下文(如患者病史和手术策略)则在摘要文本中描述。先前工作大多将这些异构信号压缩到单一平坦嵌入空间,隐含假设各层级间相互独立。但这并不理想,因为它忽略了跨层级语义包含关系(如操作属于步骤,步骤构成阶段),削弱了长距离依赖建模。为此,我们提出一种双曲空间下的手术视频-语言预训练框架,通过缓解流程上下文引发的结构假负例,并强化父阶段与其子步骤间的语义一致性,显式保留层级结构。在多个手术基准上的实验表明,该方法在不同手术类型和机构间实现了零样本与少样本阶段识别的一致性提升。

原文摘要 · Abstract (English)

Surgical vision-language foundation models typically adopt educational materials, such as surgical lecture videos, to transfer surgical knowledge encoded in language into visual representations. These knowledge are multi-dimensional and hierarchical: fine-grained action cues appear in narration, mid-level key steps are summarized in subsection headings, and global procedural context, such as patient history and surgical strategy, is described in abstract texts. Prior work largely collapses these heterogeneous signals into a single flat embedding space, implicitly assuming independence across hierarchy levels. However, this is suboptimal because it ignores cross-level semantic containment, e.g., actions belong to steps, steps compose phases, weakens long-range dependency modeling. To this end, we propose a hyperbolic surgical video-language pre-training framework that explicitly preserves the hierarchical structure by mitigating structural false negatives induced by procedural context and enforcing semantic consistency between parent phases and their constituent child steps. Extensive experiments on multiple surgical benchmarks show consistent gains in zero- and few-shot phase recognition across procedures and institutions.

手术视觉双曲空间多模态预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。