用文本提示提升视频动作定位精度,无需标注数据
Zero-Shot Temporal Action Localization Through Textual Guidance

- 利用大模型生成的文本信息作为动作定位的细粒度引导
- 在THUMOS14和ActivityNet-v1.3上超越现有无监督方法
- 适合关注零样本视频理解与少数据场景的研究者
零样本时序动作定位(ZS-TAL)旨在对未见动作类别进行分类与定位,且训练时无对应标签。现有方法依赖视觉语言模型(VLMs),但其细粒度分类能力不足,难以区分动作存在与否。多数方法通过大规模带标注视频训练,受限于数据需求且泛化能力有限。本文提出无需标注数据的新方法TEGU,通过利用大语言模型生成的丰富文本信息及字幕中提取的结构化文本,增强对视频中细微动作差异的判别能力。实验在THUMOS14和ActivityNet-v1.3数据集上验证,TEGU在不使用训练标签的情况下,优于当前先进无监督方法。
原文摘要 · Abstract (English)
Zero-shot temporal action localization (ZS-TAL) consists of classifying and localizing actions in untrimmed videos, where action classes are unseen at training time. Existing work uses Vision and Language Models (VLMs), taking advantage of their strong zero-shot transfer capabilities. Yet, these models face evident challenges with fine-grained action classification, making it difficult to directly use them to distinguish between the presence and absence of an action. Most current methods for ZS-TAL address these challenges by training models on large-scale video datasets, which require annotated data and often result in limited generalization performance. Recently, approaches discarding the use of labeled data have emerged as an alternative. Following this direction, we propose a novel approach, ``Textual Guidance for finer localization of actions in videos'' (TEGU), that compensates for the lack of supervision from training data by exploiting rich textual information derived from large language models and structured text extracted from captions. This additional linguistic context can improve fine-grained discrimination by providing richer cues about fine-grained action differences within videos. We validate the effectiveness of the proposed method by conducting experiments on the THUMOS14 and the ActivityNet-v1.3 datasets. Our results show that, by exploiting rich textual information for improved action localization, TEGU outperforms state-of-the-art ZS-TAL approaches that do not involve training
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。