用视觉语言模型提升手术流程分析的特征表示能力。
ReSW-VL: Representation Learning for Surgical Workflow Analysis Using Vision-Language Model
- 基于CLIP模型,通过提示学习微调图像编码器。
- 在三个数据集上显著优于传统方法。
- 适合医疗影像分析与智能手术辅助研究者。
从视频中识别手术阶段是一项自动分类手术进程的技术,具有实时手术支持、医疗资源优化、培训与技能评估及安全提升等广泛应用前景。近年来,手术阶段识别技术主要聚焦于基于Transformer的方法,尽管利用CNN提取单帧空间特征,并通过时间序列建模处理特征序列的方法也表现出高性能。然而,针对用于特征提取的CNN训练方法或表示学习的研究仍显不足。本研究提出一种基于视觉语言模型的手术流程分析表征学习方法(ReSW-VL)。该方法通过提示学习对CLIP视觉语言模型的图像编码器进行微调,以适应手术阶段识别任务。在三个手术阶段识别数据集上的实验结果表明,所提方法在性能上显著优于传统方法。
原文摘要 · Abstract (English)
Surgical phase recognition from video is a technology that automatically classifies the progress of a surgical procedure and has a wide range of potential applications, including real-time surgical support, optimization of medical resources, training and skill assessment, and safety improvement. Recent advances in surgical phase recognition technology have focused primarily on Transform-based methods, although methods that extract spatial features from individual frames using a CNN and video features from the resulting time series of spatial features using time series modeling have shown high performance. However, there remains a paucity of research on training methods for CNNs employed for feature extraction or representation learning in surgical phase recognition. In this study, we propose a method for representation learning in surgical workflow analysis using a vision-language model (ReSW-VL). Our proposed method involves fine-tuning the image encoder of a CLIP (Convolutional Language Image Model) vision-language model using prompt learning for surgical phase recognition. The experimental results on three surgical phase recognition datasets demonstrate the effectiveness of the proposed method in comparison to conventional methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。