无需训练即可精确定位视频中的动作,效果超越现有方法。
VideoGEM: Training-free Action Grounding in Videos

- 利用高层语义特征,通过动态加权注意力层提升动作定位能力。
- 在四个数据集上优于已有训练过的顶尖模型,最高提升12.3%。
- 适合研究零样本视频理解、动作定位的学者和工程师。
视觉语言基础模型在零样本任务中表现出色,尤其在图像中目标定位方面。然而,将这些能力用于视频中的动作与事件定位仍具挑战性,因动作缺乏明显物理轮廓且常由高层次概念描述。本文提出VideoGEM,首个基于预训练图像-视频语言模型的无训练空间动作定位方法。通过将GEM的自注意力机制拓展至视频场景,我们发现高层语义概念(如动作)通常出现在模型深层。为此,设计了层加权策略以优先考虑高层特征,并引入动态加权机制自动调节各层对特定提示的相关性。此外,采用提示分解策略,分别处理动作、动词和对象提示,显著提升动作空间定位精度。我们在CLIP、OpenCLIP和ViCLIP三个模型上评估,涵盖V-HICO、DALY、YouCook-Interactions和GroundingYouTube四个视频定位数据集,结果表明该方法无需训练即可超越当前最优有监督模型,在多个指标上实现12.3%的绝对提升。
原文摘要 · Abstract (English)
Vision-language foundation models have shown impressive capabilities across various zero-shot tasks, including training-free localization and grounding, primarily focusing on localizing objects in images. However, leveraging those capabilities to localize actions and events in videos is challenging, as actions have less physical outline and are usually described by higher-level concepts. In this work, we propose VideoGEM, the first training-free spatial action grounding method based on pretrained image- and video-language backbones. Namely, we adapt the self-self attention formulation of GEM to spatial activity grounding. We observe that high-level semantic concepts, such as actions, usually emerge in the higher layers of the image- and video-language models. We, therefore, propose a layer weighting in the self-attention path to prioritize higher layers. Additionally, we introduce a dynamic weighting method to automatically tune layer weights to capture each layer`s relevance to a specific prompt. Finally, we introduce a prompt decomposition, processing action, verb, and object prompts separately, resulting in a better spatial localization of actions. We evaluate the proposed approach on three image- and video-language backbones, CLIP, OpenCLIP, and ViCLIP, and on four video grounding datasets, V-HICO, DALY, YouCook-Interactions, and GroundingYouTube, showing that the proposed training-free approach is able to outperform current trained state-of-the-art approaches for spatial video grounding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。