用手术视频训练多模态模型,提升术中事件识别精度
CliPPER: Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition
- 设计上下文感知的视频文本对比学习与片段顺序预测预训练策略
- 在多个公开手术数据集上实现零样本事件识别新纪录
- 适合医疗视觉理解、智能手术辅助等研究者参考
视频-语言基础模型在众多任务中展现强大零样本能力。然而,在术中手术过程领域,标注数据稀缺且对时间细节要求高,挑战巨大。为此,我们提出CliPPER(长时手术视频上下文化视频-语言预训练框架),基于手术讲座视频进行训练,旨在实现细粒度的时间视频-文本识别。方法引入上下文视频-文本对比学习(VTC_CTX)和片段顺序预测(COP)两项新预训练目标,利用时间与上下文依赖增强局部视频理解;同时采用同一手术视频内视频-文本匹配的循环一致性对齐机制,强化双向一致性并提升表征连贯性;还引入帧-文本匹配(FTM)损失,进一步优化帧与文本间对齐。实验表明,该模型在多个公开手术基准上均取得新最佳性能,涵盖零样本阶段、步骤、器械及三元组识别。代码与预训练文本见https://github.com/CAMMA-public/CliPPER。
原文摘要 · Abstract (English)
Video-language foundation models have proven to be highly effective in zero-shot applications across a wide range of tasks. A particularly challenging area is the intraoperative surgical procedure domain, where labeled data is scarce, and precise temporal understanding is often required for complex downstream tasks. To address this challenge, we introduce CliPPER (Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition), a novel video-language pretraining framework trained on surgical lecture videos. Our method is designed for fine-grained temporal video-text recognition and introduces several novel pretraining strategies to improve multimodal alignment in long-form surgical videos. Specifically, we propose Contextual Video-Text Contrastive Learning (VTC_CTX) and Clip Order Prediction (COP) pretraining objectives, both of which leverage temporal and contextual dependencies to enhance local video understanding. In addition, we incorporate a Cycle-Consistency Alignment over video-text matches within the same surgical video to enforce bidirectional consistency and improve overall representation coherence. Moreover, we introduce a more refined alignment loss, Frame-Text Matching (FTM), to improve the alignment between video frames and text. As a result, our model establishes a new state-of-the-art across multiple public surgical benchmarks, including zero-shot recognition of phases, steps, instruments, and triplets. The source code and pretraining captions can be found at https://github.com/CAMMA-public/CliPPER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。