arXiv:2410.14238cs.CV2024-10NeurIPS被引 2

用分镜式细粒度描述提升视频动作识别精度

Storyboard guided Alignment for Fine-grained Video Action Recognition

  • 基于分镜思想生成视频原子动作描述,增强语义对齐
  • 在多个数据集上实现监督/少样本/零样本下的最优性能
  • 适合需要精准理解动作细节的视频分析任务

细粒度视频动作识别可视为视频-文本匹配问题。以往方法依赖全局视频语义整合视频嵌入,因缺乏对原子动作语义的理解,易导致视频-文本对齐偏差。本文提出多粒度框架,基于两点观察:(i) 不同全局语义的视频可能共享相似原子动作或外观;(ii) 视频中的原子动作可能是瞬时、缓慢或与全局语义无直接关联。受分镜概念启发,利用预训练大语言模型生成细粒度视频描述,捕捉视频中常见的原子动作。设计过滤度量筛选出视频与描述中共同存在的原子动作描述。结合全局语义与细粒度描述,定位关键帧并聚合嵌入,提升嵌入准确性。在多个视频动作识别数据集上的大量实验表明,该方法在监督、少样本和零样本设置下均表现优异。

原文摘要 · Abstract (English)

Fine-grained video action recognition can be conceptualized as a video-text matching problem. Previous approaches often rely on global video semantics to consolidate video embeddings, which can lead to misalignment in video-text pairs due to a lack of understanding of action semantics at an atomic granularity level. To tackle this challenge, we propose a multi-granularity framework based on two observations: (i) videos with different global semantics may share similar atomic actions or appearances, and (ii) atomic actions within a video can be momentary, slow, or even non-directly related to the global video semantics. Inspired by the concept of storyboarding, which disassembles a script into individual shots, we enhance global video semantics by generating fine-grained descriptions using a pre-trained large language model. These detailed descriptions capture common atomic actions depicted in videos. A filtering metric is proposed to select the descriptions that correspond to the atomic actions present in both the videos and the descriptions. By employing global semantics and fine-grained descriptions, we can identify key frames in videos and utilize them to aggregate embeddings, thereby making the embedding more accurate. Extensive experiments on various video action recognition datasets demonstrate superior performance of our proposed method in supervised, few-shot, and zero-shot settings.

动作识别细粒度视频理解分镜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。