用专家视频和语言理解提升手术视频分析,无需大量标注数据。
Watch and Learn: Leveraging Expert Knowledge and Language for Surgical Video Understanding
- 模仿人类看专家视频学知识,用视频-语言模型学习时空特征。
- 在两个手术领域实现最高7%的阶段分割提升,零样本任务增8%。
- 适合医疗AI研究者,可做手术视频密集描述生成等新任务。
自动化手术流程分析对教学、研究和临床决策至关重要,但缺乏标注数据限制了准确全面的分析方法发展。我们提出一种受人类学习过程启发的新方法:通过观看专家视频并理解其解释来学习。该方法利用在对齐、去噪和生成任务上训练的视频-语言模型,学习短期时空与多模态表示,并使用特定任务的时间模型捕捉全视频关系。为实现手术领域的全面视频-语言理解,我们设计了一种数据收集与过滤策略,从教育类YouTube视频构建大规模预训练数据集。随后采用参数高效微调,将公开手术数据集的任务标注映射到语言域。在两个手术领域上的实验表明,本方法在阶段分割任务中性能提升达7%,零样本阶段分割提升8%,少样本设置下表现媲美全监督模型。利用模型的长时序定位与文本生成能力,我们首次实现手术视频的密集视频描述(DVC),解决了该领域无现成数据集的问题。本方法融合视频-语言预训练、大规模视频预训练与优化微调,优于现有技术并开启新下游任务。
原文摘要 · Abstract (English)
Automated surgical workflow analysis is crucial for education, research, and clinical decision-making, but the lack of annotated datasets hinders the development of accurate and comprehensive workflow analysis solutions. We introduce a novel approach for addressing the sparsity and heterogeneity of annotated training data inspired by the human learning procedure of watching experts and understanding their explanations. Our method leverages a video-language model trained on alignment, denoising, and generative tasks to learn short-term spatio-temporal and multimodal representations. A task-specific temporal model is then used to capture relationships across entire videos. To achieve comprehensive video-language understanding in the surgical domain, we introduce a data collection and filtering strategy to construct a large-scale pretraining dataset from educational YouTube videos. We then utilize parameter-efficient fine-tuning by projecting downstream task annotations from publicly available surgical datasets into the language domain. Extensive experiments in two surgical domains demonstrate the effectiveness of our approach, with performance improvements of up to 7% in phase segmentation tasks, 8% in zero-shot phase segmentation, and comparable capabilities to fully-supervised models in few-shot settings. Harnessing our model's capabilities for long-range temporal localization and text generation, we present the first comprehensive solution for dense video captioning (DVC) of surgical videos, addressing this task despite the absence of existing DVC datasets in the surgical domain. We introduce a novel approach to surgical workflow understanding that leverages video-language pretraining, large-scale video pretraining, and optimized fine-tuning. Our method improves performance over state-of-the-art techniques and enables new downstream tasks for surgical video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。