arXiv:2409.20083cs.CV2024-09被引 5

用少量参数将图像模型高效迁移至手术视频阶段识别,提升精度与泛化能力。

SurgPETL: Parameter-Efficient Image-to-Surgical-Video Transfer Learning for Surgical Phase Recognition

论文配图:SurgPETL: Parameter-Efficient Image-to-Surgical-Video Transfer Learning for Surgical Phase Recognition
图 1 · 摘自论文原文
  • 基于视觉变压器构建参数高效的跨模态迁移框架,融合空间与时间适配器。
  • 在三个手术数据集上达到新最优,比基线提升超过5%准确率。
  • 适合医疗视频分析、少样本学习及资源受限场景的研究者使用。

利用图像预训练模型在下游任务中展现出优异性能,但“图像预训练后视频微调”的范式在高维视频数据上存在显著性能瓶颈。尤其在医学领域,手术视频任务受限于数据稀缺性且需全面建模时空特征。近期,参数高效的图像到视频迁移学习成为视频动作识别的有效方法,利用具备强特征迁移能力的图像模型,并通过最小微调实现跨模态时间建模。然而,该范式在复杂手术领域的有效性与通用性尚未探索。本文提出一种新型问题:高效适应图像预训练模型以专用于细粒度手术阶段识别,即参数高效的图像到手术视频迁移学习(SurgPETL)。首先,构建了面向手术阶段识别的参数高效迁移学习基准SurgPETL,基于两种不同规模的ViT,在五个大规模自然与医学数据集上预训练,评估三种先进方法。其次,提出空间-时间适配模块(STA),结合标准空间适配器与新型时间适配器,捕获精细空间特征并建立时序连接,实现鲁棒的时空建模。在涵盖多种手术流程的三个挑战性数据集上的实验表明,SurgPETL配合STA显著有效。

原文摘要 · Abstract (English)

Capitalizing on image-level pre-trained models for various downstream tasks has recently emerged with promising performance. However, the paradigm of "image pre-training followed by video fine-tuning" for high-dimensional video data inevitably poses significant performance bottlenecks. Furthermore, in the medical domain, many surgical video tasks encounter additional challenges posed by the limited availability of video data and the necessity for comprehensive spatial-temporal modeling. Recently, Parameter-Efficient Image-to-Video Transfer Learning has emerged as an efficient and effective paradigm for video action recognition tasks, which employs image-level pre-trained models with promising feature transferability and involves cross-modality temporal modeling with minimal fine-tuning. Nevertheless, the effectiveness and generalizability of this paradigm within intricate surgical domain remain unexplored. In this paper, we delve into a novel problem of efficiently adapting image-level pre-trained models to specialize in fine-grained surgical phase recognition, termed as Parameter-Efficient Image-to-Surgical-Video Transfer Learning. Firstly, we develop a parameter-efficient transfer learning benchmark SurgPETL for surgical phase recognition, and conduct extensive experiments with three advanced methods based on ViTs of two distinct scales pre-trained on five large-scale natural and medical datasets. Then, we introduce the Spatial-Temporal Adaptation module, integrating a standard spatial adapter with a novel temporal adapter to capture detailed spatial features and establish connections across temporal sequences for robust spatial-temporal modeling. Extensive experiments on three challenging datasets spanning various surgical procedures demonstrate the effectiveness of SurgPETL with STA.

手术视频迁移学习高效微调时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。