用视觉大模型提升小切口白内障手术阶段分割,标签少也能高效准确。
Data-Efficient Surgical Phase Segmentation in Small-Incision Cataract Surgery: A Controlled Study of Vision Foundation Models

- 统一时间模型,对比自监督大模型与传统网络的视觉表征能力
- DINOv3 ViT-7B达83.4%准确率、87.0编辑分数,最优表现
- 提供低标注医疗视频场景下的实用迁移学习方案,适合医学影像研究者
手术阶段分割是计算机辅助手术的核心,但在标注手术视频稀缺时仍难构建鲁棒模型。本研究通过受控实验,探讨手动小切口白内障手术(SICS)中的数据高效阶段分割。为分离表征质量影响,所有视觉编码器在相同时间模型(MS-TCN++)和训练评估设置下,于SICS-155数据集(19个阶段)上进行对比。比较了监督编码器(ResNet-50, I3D)与大型自监督基础模型(DINOv3, V-JEPA2),并采用缓存特征流水线,将昂贵的视觉编码与轻量时间学习解耦。结果显示,基础模型特征显著提升分割性能,其中DINOv3 ViT-7B达到最高整体效果:83.4%准确率,87.0编辑分数。进一步研究无标签视频上的白内障领域迁移与轻量适应,分析其增益或损害条件。总体表明,现代视觉基础模型对手术流程理解具有强迁移性,并为低标签医疗视频场景提供实用指导。
原文摘要 · Abstract (English)
Surgical phase segmentation is central to computer-assisted surgery, yet robust models remain difficult to develop when labeled surgical videos are scarce. We study data-efficient phase segmentation for manual small-incision cataract surgery (SICS) through a controlled comparison of visual representations. To isolate representation quality, we pair each visual encoder with the same temporal model (MS-TCN++) under identical training and evaluation settings on SICS-155 (19 phases). We compare supervised encoders (ResNet-50, I3D) against large self-supervised foundation models (DINOv3, V-JEPA2), and use a cached-feature pipeline that decouples expensive visual encoding from lightweight temporal learning. Foundation-model features improve segmentation performance in this setup, with DINOv3 ViT-7B achieving the best overall results (83.4% accuracy, 87.0 edit score). We further examine cataract-domain transfer using unlabeled videos and lightweight adaptation, and analyze when it helps or hurts. Overall, the study indicates strong transferability of modern vision foundation models to surgical workflow understanding and provides practical guidance for low-label medical video settings. The project website is available at: https://sl2005.github.io/DataEfficient-sics-phase-seg/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。