arXiv:2506.02692cs.CV2025-06被引 24

首个面向手术视频的时空联合预训练模型,提升手术动态理解能力。

Large-scale Self-supervised Video Foundation Model for Intelligent Surgery

  • 通过联合建模空间与时间特征,实现手术视频的端到端预训练。
  • 在13个手术数据集上优于自然与手术领域预训练模型,显著提升理解能力。
  • 适合智能手术辅助系统研发者、医学影像算法工程师参考使用。

计算机辅助干预有望革新现代外科手术,其中手术场景理解是支持决策、提高操作效率和保障术中安全的关键。现有基于AI的方法虽通过自监督空间表征学习减轻标注负担,但预训练阶段缺乏显式的时间建模,限制了对动态手术上下文的捕捉,导致时空理解不完整。本文提出首个面向手术视频的视频级预训练框架,实现大规模手术视频数据上的联合时空表征学习。我们构建了包含3,650段视频、约355万帧的大型手术视频数据集,覆盖20余种手术操作和10余个解剖结构。基于此,提出SurgVISTA(手术视频级时空架构)——一种基于重建的预训练方法,通过联合时空建模捕捉复杂的空间结构与时间动态;同时引入手术专家指导的图像级知识蒸馏,增强对细微解剖与语义特征的学习。为验证效果,建立了涵盖六个手术类型、四个任务的13个视频级基准数据集。大量实验表明,SurgVISTA在各项任务中均优于自然域及手术域预训练模型,展现出在临床有意义场景中推动智能手术系统发展的强大潜力。

原文摘要 · Abstract (English)

Computer-Assisted Intervention (CAI) has the potential to revolutionize modern surgery, with surgical scene understanding serving as a critical component in supporting decision-making, improving procedural efficacy, and ensuring intraoperative safety. While existing AI-driven approaches alleviate annotation burdens via self-supervised spatial representation learning, their lack of explicit temporal modeling during pre-training fundamentally restricts the capture of dynamic surgical contexts, resulting in incomplete spatiotemporal understanding. In this work, we introduce the first video-level surgical pre-training framework that enables joint spatiotemporal representation learning from large-scale surgical video data. To achieve this, we constructed a large-scale surgical video dataset comprising 3,650 videos and approximately 3.55 million frames, spanning more than 20 surgical procedures and over 10 anatomical structures. Building upon this dataset, we propose SurgVISTA (Surgical Video-level Spatial-Temporal Architecture), a reconstruction-based pre-training method that captures intricate spatial structures and temporal dynamics through joint spatiotemporal modeling. Additionally, SurgVISTA incorporates image-level knowledge distillation guided by a surgery-specific expert to enhance the learning of fine-grained anatomical and semantic features. To validate its effectiveness, we established a comprehensive benchmark comprising 13 video-level datasets spanning six surgical procedures across four tasks. Extensive experiments demonstrate that SurgVISTA consistently outperforms both natural- and surgical-domain pre-trained models, demonstrating strong potential to advance intelligent surgical systems in clinically meaningful scenarios.

手术视觉视频预训练时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。