arXiv:2510.27255cs.CV2025-10被引 3

用语言描述属性提升视频动作零样本识别准确率

Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes

  • 从网络抓取描述,用大模型提取关键词替代人工标注
  • 在UCF-101等数据集上达81.0%、53.1%、68.9%准确率
  • 适合做零样本动作识别且减少人工标注成本

视觉语言模型(VLMs)在零样本动作识别中表现出色,通过学习视频嵌入与类别嵌入的关联。然而,仅依赖动作类别提供语义上下文时,易受多义词干扰,导致理解模糊。为此,本文提出一种新方法:利用大规模语言模型从网络爬取的描述中自动提取相关关键词,降低对人工标注的依赖,避免繁琐的人工属性构建过程。此外,设计了时空交互模块,聚焦物体与动作单元,促进描述属性与视频内容的对齐。在零样本实验中,模型在UCF-101、HMDB-51和Kinetics-600上分别达到81.0%、53.1%和68.9%的准确率,证明其在多种下游任务中的适应性与有效性。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated impressive capabilities in zero-shot action recognition by learning to associate video embeddings with class embeddings. However, a significant challenge arises when relying solely on action classes to provide semantic context, particularly due to the presence of multi-semantic words, which can introduce ambiguity in understanding the intended concepts of actions. To address this issue, we propose an innovative approach that harnesses web-crawled descriptions, leveraging a large-language model to extract relevant keywords. This method reduces the need for human annotators and eliminates the laborious manual process of attribute data creation. Additionally, we introduce a spatio-temporal interaction module designed to focus on objects and action units, facilitating alignment between description attributes and video content. In our zero-shot experiments, our model achieves impressive results, attaining accuracies of 81.0%, 53.1%, and 68.9% on UCF-101, HMDB-51, and Kinetics-600, respectively, underscoring the model's adaptability and effectiveness across various downstream tasks.

零样本识别视觉语言模型动作识别自动标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。