arXiv:2410.00263cs.CVcs.AI2024-10NeurIPS被引 54

用分层知识增强提升手术视频与语言的对齐能力。

Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation

  • 用大模型增强手术概念,弥补文字信息丢失。
  • 结合语言与视觉自监督,提升跨模态对齐效果。
  • 在多个数据集上实现零样本迁移领先性能。

手术视频-语言预训练(VLP)因领域知识鸿沟和多模态数据稀缺面临挑战。本文提出分层知识增强方法与新型程序编码手术知识增强视频-语言预训练框架(PeskaVLP),以解决手术讲座视频中的文本信息丢失及时空对齐难题。通过大语言模型(LLM)精炼并丰富手术概念,提供全面的语言监督,降低过拟合风险。PeskaVLP融合语言监督与视觉自监督,构建硬负样本,并采用基于动态时间规整(DTW)的损失函数,有效理解跨模态流程对齐。在多个公开的手术场景理解与跨模态检索数据集上进行的大量实验表明,该方法显著提升零样本迁移性能,为手术场景理解提供通用视觉表征。代码已开源:https://github.com/CAMMA-public/SurgVLP。

原文摘要 · Abstract (English)

Surgical video-language pretraining (VLP) faces unique challenges due to the knowledge domain gap and the scarcity of multi-modal data. This study aims to bridge the gap by addressing issues regarding textual information loss in surgical lecture videos and the spatial-temporal challenges of surgical VLP. We propose a hierarchical knowledge augmentation approach and a novel Procedure-Encoded Surgical Knowledge-Augmented Video-Language Pretraining (PeskaVLP) framework to tackle these issues. The knowledge augmentation uses large language models (LLM) for refining and enriching surgical concepts, thus providing comprehensive language supervision and reducing the risk of overfitting. PeskaVLP combines language supervision with visual self-supervision, constructing hard negative samples and employing a Dynamic Time Warping (DTW) based loss function to effectively comprehend the cross-modal procedural alignment. Extensive experiments on multiple public surgical scene understanding and cross-modal retrieval datasets show that our proposed method significantly improves zero-shot transferring performance and offers a generalist visual representation for further advancements in surgical scene understanding.The code is available at https://github.com/CAMMA-public/SurgVLP

视频语言预训练手术视觉多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。