arXiv:2502.03549cs.CV2025-02ICLR被引 7

让CLIP学会理解视频中的动作行为,提升视频理解能力。

Kronecker Mask and Interpretive Prompts are Language-Action Video Learners

  • 用克罗内克掩码增强时间建模,扩大感受野并保持时空异质性。
  • 通过大语言模型生成语义丰富的动作描述提示,强化动词理解。
  • 适用于多种视频任务,对动作识别有显著提升,适合视频理解研究者。

对比语言图像预训练(CLIP)显著推动了基于图像的视觉学习发展。一个紧迫的问题随之出现:如何有效将CLIP适配到视频领域?现有研究主要调整CLIP的文本或视觉分支以实现动作识别,但我们认为两者都需改进。本文提出 extbf{CLAVER}:一种对比语言-动作-视频学习器,旨在将CLIP的关注点从静态视觉对象与具体名词的对齐,转向动态动作行为与抽象动词的对齐。我们引入一种新颖的克罗内克掩码注意力机制用于时间建模,该机制具备三大优势:1)扩展每个标记的时间感受野;2)作为有效的时空异质性归纳偏置,缓解时空同质化问题;3)可无缝嵌入基于Transformer的模型中。在文本分支方面,利用大语言模型生成多样、句子级且语义丰富的动作解释性提示,引导模型聚焦于动词理解。在多个基准和学习场景下的大量实验表明,本方法具有优越性和通用性。

原文摘要 · Abstract (English)

Contrastive language-image pretraining (CLIP) has significantly advanced image-based vision learning. A pressing topic subsequently arises: how can we effectively adapt CLIP to the video domain? Recent studies have focused on adjusting either the textual or visual branch of CLIP for action recognition. However, we argue that adaptations of both branches are crucial. In this paper, we propose \textbf{CLAVER}: a \textbf{C}ontrastive \textbf{L}anguage-\textbf{A}ction \textbf{V}ideo Learn\textbf{er}, designed to shift CLIP's focus from the alignment of static visual objects and concrete nouns to the alignment of dynamic action behaviors and abstract verbs. Specifically, we introduce a novel Kronecker mask attention for temporal modeling. Our tailored Kronecker mask offers three benefits 1) it expands the temporal receptive field for each token, 2) it serves as an effective spatiotemporal heterogeneity inductive bias, mitigating the issue of spatiotemporal homogenization, and 3) it can be seamlessly plugged into transformer-based models. Regarding the textual branch, we leverage large language models to generate diverse, sentence-level and semantically rich interpretive prompts of actions, which shift the model's focus towards the verb comprehension. Extensive experiments on various benchmarks and learning scenarios demonstrate the superiority and generality of our approach.

视频理解CLIP动作识别提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。