arXiv:2502.08822cs.CV2025-02中稿 · IEEE International…

基于重要性掩码的预训练模型,提升白内障手术视频分析效果

$\mathsf{CSMAE~}$:~Cataract Surgical Masked Autoencoder (MAE) based Pre-training

  • 按时空重要性选择掩码位置,优化自编码器学习
  • 在两个数据集上超越现有自监督方法,显著提升步骤识别准确率
  • 适合需要高效手术视频分析的研究者与医疗AI开发者

自动化分析手术视频对提升外科培训、流程优化和术后评估至关重要。本文提出一种专为白内障手术视频设计的掩码自编码器(CSMAE)预训练方法,不同于随机掩码,该方法依据视频片段的时空重要性进行掩码选择。我们构建了一个大规模白内障手术视频数据集,以提高模型的学习效率并增强其在低数据场景下的鲁棒性。预训练模型可通过微调快速适配具体下游任务,作为后续分析的可靠骨干网络。在两个白内障手术视频数据集D99和Cataract-101上的步骤识别任务测试中,本方法显著优于当前最先进的自监督预训练及适配器迁移学习方法,不仅展示了MAE在手术视频分析中的潜力,也树立了新的研究基准。

原文摘要 · Abstract (English)

Automated analysis of surgical videos is crucial for improving surgical training, workflow optimization, and postoperative assessment. We introduce a CSMAE, Masked Autoencoder (MAE)-based pretraining approach, specifically developed for Cataract Surgery video analysis, where instead of randomly selecting tokens for masking, they are selected based on the spatiotemporal importance of the token. We created a large dataset of cataract surgery videos to improve the model's learning efficiency and expand its robustness in low-data regimes. Our pre-trained model can be easily adapted for specific downstream tasks via fine-tuning, serving as a robust backbone for further analysis. Through rigorous testing on a downstream step-recognition task on two Cataract Surgery video datasets, D99 and Cataract-101, our approach surpasses current state-of-the-art self-supervised pretraining and adapter-based transfer learning methods by a significant margin. This advancement not only demonstrates the potential of our MAE-based pretraining in the field of surgical video analysis but also sets a new benchmark for future research.

手术视频自监督学习掩码自编码器白内障手术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。