arXiv:2602.13633cs.CV2026-02

ZEN模型让手术视频理解跨术式通用,提升AI辅助诊疗能力

A generalizable foundation model for intraoperative understanding across surgical procedures

  • 用自监督多教师蒸馏法训练,覆盖21种手术、超400万帧视频
  • 在20项任务中均优于现有模型,零样本下仍保持强泛化性能
  • 适合做手术辅助系统和标准化培训评估的通用视觉基础模型

微创手术中,临床决策依赖实时视觉判断,但不同外科医生和手术类型间存在显著感知差异,限制了评估一致性、培训标准化及可靠AI系统的开发。现有手术AI模型多针对特定任务,难以跨术式或机构泛化。本文提出ZEN——一种可泛化的手术视频理解基础模型,基于超过21种手术、逾400万帧视频,采用自监督多教师蒸馏框架进行训练。研究构建了大规模多样化数据集,并在统一基准下系统评估多种表示学习策略。在20个下游任务中,无论全微调、冻结主干、少样本或零样本设置,ZEN均显著优于现有手术基础模型,展现出强大的跨术式泛化能力。结果表明,该模型推动了手术场景理解的统一表征发展,为术中辅助与手术培训评估提供新可能。

原文摘要 · Abstract (English)

In minimally invasive surgery, clinical decisions depend on real-time visual interpretation, yet intraoperative perception varies substantially across surgeons and procedures. This variability limits consistent assessment, training, and the development of reliable artificial intelligence systems, as most surgical AI models are designed for narrowly defined tasks and do not generalize across procedures or institutions. Here we introduce ZEN, a generalizable foundation model for intraoperative surgical video understanding trained on more than 4 million frames from over 21 procedures using a self-supervised multi-teacher distillation framework. We curated a large and diverse dataset and systematically evaluated multiple representation learning strategies within a unified benchmark. Across 20 downstream tasks and full fine-tuning, frozen-backbone, few-shot and zero-shot settings, ZEN consistently outperforms existing surgical foundation models and demonstrates robust cross-procedure generalization. These results suggest a step toward unified representations for surgical scene understanding and support future applications in intraoperative assistance and surgical training assessment.

手术视觉基础模型泛化能力自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。