VidLPRO提升手术视频理解,零样本识别精度超现有模型21.5%。
VidLPRO: A $\underline{Vid}$eo-$\underline{L}$anguage $\underline{P}$re-training Framework for $\underline{Ro}$botic and Laparoscopic Surgery
- 融合三重学习目标,捕捉手术视频与语言的时序对齐关系。
- 在Cholec80和AutoLaparo上实现最高21.5%准确率提升。
- 支持单帧推理且可扩展,适合临床部署与研究应用。
我们提出VidLPRO,一种专为机器人与腹腔镜手术设计的视频-语言(VL)预训练框架。现有手术VL模型多依赖对比学习,而本工作采用更全面的方法,以捕捉复杂的时序动态并实现视频与语言的精准对齐。VidLPRO整合了视频-文本对比学习、视频-文本匹配及掩码语言建模三项目标,以学习丰富的多模态表征。为此,我们构建了GenSurg+数据集,基于GenSurgery衍生而来,包含17,000个手术视频片段,其字幕由Whisper提取转录后,经GPT-4生成。该数据集解决了手术领域高质量大规模VL数据稀缺的问题。在Cholec80与AutoLaparo等基准上的大量实验表明,VidLPRO在零样本手术阶段识别任务中达到当前最优性能,显著超越SurgVLP与HecVL等现有模型,准确率提升最高达21.5%,F1分数提升15.7%。值得注意的是,该模型在单帧输入下仍保持稳健表现,并能随上下文时间范围扩大有效扩展。消融研究揭示了帧采样策略对性能与效率的影响。这些结果证明VidLPRO具备作为手术视频理解基础模型的巨大潜力。
原文摘要 · Abstract (English)
We introduce VidLPRO, a novel video-language (VL) pre-training framework designed specifically for robotic and laparoscopic surgery. While existing surgical VL models primarily rely on contrastive learning, we propose a more comprehensive approach to capture the intricate temporal dynamics and align video with language. VidLPRO integrates video-text contrastive learning, video-text matching, and masked language modeling objectives to learn rich VL representations. To support this framework, we present GenSurg+, a carefully curated dataset derived from GenSurgery, comprising 17k surgical video clips paired with captions generated by GPT-4 using transcripts extracted by the Whisper model. This dataset addresses the need for large-scale, high-quality VL data in the surgical domain. Extensive experiments on benchmark datasets, including Cholec80 and AutoLaparo, demonstrate the efficacy of our approach. VidLPRO achieves state-of-the-art performance in zero-shot surgical phase recognition, significantly outperforming existing surgical VL models such as SurgVLP and HecVL. Our model demonstrates improvements of up to 21.5\% in accuracy and 15.7% in F1 score, setting a new benchmark in the field. Notably, VidLPRO exhibits robust performance even with single-frame inference, while effectively scaling with increased temporal context. Ablation studies reveal the impact of frame sampling strategies on model performance and computational efficiency. These results underscore VidLPRO's potential as a foundation model for surgical video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。