arXiv:2412.09895cs.CV2024-12AAAI被引 16

用双模态时空动态机制提升CLIP零样本动作识别能力

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP

  • 在视觉侧设计时空交叉注意力,无额外参数捕捉动态特征
  • 构建动作语义知识图谱生成细粒度文本提示,提升语义理解
  • 无需微调即可在三个基准上超越现有方法,适合零样本场景

零样本动作识别(ZSAR)依赖多模态时空理解。然而,直接微调CLIP在捕捉视觉与文本的时序动态方面表现不佳,尤其面对具有细微时空差异的新动作时。本文提出基于CLIP的新型框架STDD,协同建模多模态时空动态。视觉侧引入高效时空交叉注意力,在空间注意力前后施加轻量操作,灵活捕捉时空特征,不增加参数或计算开销。语义侧通过构建动作语义知识图(ASKG),综合分解动作为静态外观与动态运动,生成细粒度文本提示。训练阶段,帧级视频表示与提示级语义表示精准对齐,同时受冻结CLIP的视频表示约束以增强泛化性。大量实验表明,该方法在Kinetics-600、UCF101和HMDB51三个主流视频数据集上,于严苛的ZSAR设置下持续优于现有最优方法。

原文摘要 · Abstract (English)

Zero-shot action recognition (ZSAR) requires collaborative multi-modal spatiotemporal understanding. However, finetuning CLIP directly for ZSAR yields suboptimal performance, given its inherent constraints in capturing essential temporal dynamics from both vision and text perspectives, especially when encountering novel actions with fine-grained spatiotemporal discrepancies. In this work, we propose Spatiotemporal Dynamic Duo (STDD), a novel CLIP-based framework to comprehend multi-modal spatiotemporal dynamics synergistically. For the vision side, we propose an efficient Space-time Cross Attention, which captures spatiotemporal dynamics flexibly with simple yet effective operations applied before and after spatial attention, without adding additional parameters or increasing computational complexity. For the semantic side, we conduct spatiotemporal text augmentation by comprehensively constructing an Action Semantic Knowledge Graph (ASKG) to derive nuanced text prompts. The ASKG elaborates on static and dynamic concepts and their interrelations, based on the idea of decomposing actions into spatial appearances and temporal motions. During the training phase, the frame-level video representations are meticulously aligned with prompt-level nuanced text representations, which are concurrently regulated by the video representations from the frozen CLIP to enhance generalizability. Extensive experiments validate the effectiveness of our approach, which consistently surpasses state-of-the-art approaches on popular video benchmarks (i.e., Kinetics-600, UCF101, and HMDB51) under challenging ZSAR settings.

零样本识别多模态时空建模CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。