用470万帧手术视频训练出顶尖的外科视觉模型,性能大幅超越现有方法。
Scaling up self-supervised learning for improved surgical foundation models
- 基于超大规模手术数据集自监督预训练,构建新模型SurgeNetXL
- 在6个数据集上实现语义分割、阶段识别等任务的领先精度,提升最高达12.6%
- 开源模型与超200万帧数据,助力外科视觉研究者快速复现与拓展
基础模型通过大规模预训练显著提升了计算机视觉在多样化任务上的表现,但在外科计算机视觉领域应用有限。本研究提出SurgeNetXL,一种新型外科基础模型,创下该领域的最新基准。该模型在迄今最大规模的外科数据集(超过470万帧视频)上训练,对涵盖四种外科手术和三种任务的六个数据集均表现卓越,包括语义分割、阶段识别和关键视野安全(CVS)分类。相较于现有最优的外科基础模型,SurgeNetXL在语义分割、阶段识别和CVS分类上分别实现2.4%、9.0%和12.6%的平均提升;相比最佳的ImageNet基线模型,提升分别为14.4%、4.0%和1.6%。此外,本研究揭示了扩展预训练数据集、延长训练时长及优化模型架构对提升外科视觉模型泛化能力的关键作用,为数据稀缺场景下的鲁棒性研究提供全面框架。所有模型及部分包含超过200万帧视频的SurgeNetXL数据集已公开:https://github.com/TimJaspers0801/SurgeNet。
原文摘要 · Abstract (English)
Foundation models have revolutionized computer vision by achieving vastly superior performance across diverse tasks through large-scale pretraining on extensive datasets. However, their application in surgical computer vision has been limited. This study addresses this gap by introducing SurgeNetXL, a novel surgical foundation model that sets a new benchmark in surgical computer vision. Trained on the largest reported surgical dataset to date, comprising over 4.7 million video frames, SurgeNetXL achieves consistent top-tier performance across six datasets spanning four surgical procedures and three tasks, including semantic segmentation, phase recognition, and critical view of safety (CVS) classification. Compared with the best-performing surgical foundation models, SurgeNetXL shows mean improvements of 2.4, 9.0, and 12.6 percent for semantic segmentation, phase recognition, and CVS classification, respectively. Additionally, SurgeNetXL outperforms the best-performing ImageNet-based variants by 14.4, 4.0, and 1.6 percent in the respective tasks. In addition to advancing model performance, this study provides key insights into scaling pretraining datasets, extending training durations, and optimizing model architectures specifically for surgical computer vision. These findings pave the way for improved generalizability and robustness in data-scarce scenarios, offering a comprehensive framework for future research in this domain. All models and a subset of the SurgeNetXL dataset, including over 2 million video frames, are publicly available at: https://github.com/TimJaspers0801/SurgeNet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。