构建可扩展的手术视频预训练框架,提升模型泛化能力
Scaling Video Pretraining for Surgical Foundation Models
- 设计统一预训练流程,平衡多源数据采样
- 在16个下游任务中超越现有自监督模型表现
- 适合医疗AI研究者与手术视觉系统开发者
手术视频理解对计算机辅助干预至关重要,但现有手术基础模型受限于数据规模小、术式多样性不足及评估不一致,且缺乏可复现的训练流程。本文提出SurgRec,一种可扩展且可复现的手术视频理解预训练方案,包含SurgRec-MAE和SurgRec-JEPA两个变体。我们构建了包含10,535段视频和2.145亿帧的多源大型语料库,涵盖内窥镜、腹腔镜、白内障及机器人手术。基于此,建立统一预训练流程并实现跨16个下游数据集、4类临床场景的一致数据划分与标准化评估。在大量自监督基线与视觉语言模型对比中,SurgRec在多个下游任务中持续领先。相比之下,视觉语言模型在细粒度时序识别上表现不稳定,对提示词敏感且性能差距明显。本工作为社区提供可复现、可扩展的手术视频模型基础。所有代码、模型与数据将公开发布。
原文摘要 · Abstract (English)
Surgical video understanding is essential for computer-assisted interventions, yet existing surgical foundation models remain constrained by limited data scale, procedural diversity, and inconsistent evaluation, often lacking a reproducible training pipeline. We propose SurgRec, a scalable and reproducible pretraining recipe for surgical video understanding, instantiated with two variants: SurgRec-MAE and SurgRec-JEPA. We curate a large multi-source corpus of 10,535 videos and 214.5M frames spanning endoscopy, laparoscopy, cataract, and robotic surgery. Building on this corpus, we develop a unified pretraining pipeline with balanced sampling and standardize a reproducible benchmark across 16 downstream datasets and four clinical domains with consistent data splits. Across extensive comparisons against SSL baselines and vision-language models, SurgRec consistently achieves superior performance across downstream datasets. In contrast, VLMs prove unreliable for fine-grained temporal recognition, exhibiting both performance gaps and sensitivity to prompt phrasing. Our work provides a reproducible, scalable foundation for the community to build more general surgical video models. All code, models, and data will be publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。