arXiv:2506.20254cs.CV2025-06中稿 · MICCAI 2025被引 6

轻量级框架让手术阶段识别跨机构自适应,少样本也能精准定位。

Recognizing Surgical Phases Anywhere: Few-Shot Test-time Adaptation and Task-graph Guided Refinement

  • 用少量图像和自然语言定义阶段,实现跨机构快速适配。
  • 在多机构多手术中表现超越全量标注模型,32样本即达顶尖性能。
  • 结合任务图与扩散建模,提升测试时动态适应与时间一致性。

手术流程的复杂性与多样性受手术室环境、机构规范和解剖差异影响,制约了跨机构、跨术式通用模型的发展。尽管基于大规模视觉-语言数据预训练的手术基础模型具备良好迁移潜力,但其零样本性能仍受限于领域偏移。为此,我们提出Surgical Phase Anywhere(SPA),一种轻量级框架,可对基础模型进行极小标注下的机构化适配。SPA通过少样本空间适配,将多模态嵌入对齐至特定机构的手术场景与阶段;通过扩散建模融入基于机构手术流程的任务图先验,保障时间一致性;并采用动态测试时自适应机制,利用多模态预测流间的共识,在无标签测试视频上实现自监督调整,增强分布偏移下的可靠性。用户仅需以自然语言定义阶段、标注少量图像并提供阶段转移图,即可快速部署。实验表明,SPA在多个机构和手术类型下均达到少样本手术阶段识别的最先进水平,甚至优于使用32样本全量标注的模型。代码已开源:https://github.com/CAMMA-public/SPA

原文摘要 · Abstract (English)

The complexity and diversity of surgical workflows, driven by heterogeneous operating room settings, institutional protocols, and anatomical variability, present a significant challenge in developing generalizable models for cross-institutional and cross-procedural surgical understanding. While recent surgical foundation models pretrained on large-scale vision-language data offer promising transferability, their zero-shot performance remains constrained by domain shifts, limiting their utility in unseen surgical environments. To address this, we introduce Surgical Phase Anywhere (SPA), a lightweight framework for versatile surgical workflow understanding that adapts foundation models to institutional settings with minimal annotation. SPA leverages few-shot spatial adaptation to align multi-modal embeddings with institution-specific surgical scenes and phases. It also ensures temporal consistency through diffusion modeling, which encodes task-graph priors derived from institutional procedure protocols. Finally, SPA employs dynamic test-time adaptation, exploiting the mutual agreement between multi-modal phase prediction streams to adapt the model to a given test video in a self-supervised manner, enhancing the reliability under test-time distribution shifts. SPA is a lightweight adaptation framework, allowing hospitals to rapidly customize phase recognition models by defining phases in natural language text, annotating a few images with the phase labels, and providing a task graph defining phase transitions. The experimental results show that the SPA framework achieves state-of-the-art performance in few-shot surgical phase recognition across multiple institutions and procedures, even outperforming full-shot models with 32-shot labeled data. Code is available at https://github.com/CAMMA-public/SPA

手术识别少样本学习自适应任务图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。