arXiv:2605.09976cs.CV2026-05

无需训练即可在线识别未见过的动作,突破传统方法泛化局限。

OZ-TAL: Online Zero-Shot Temporal Action Localization

论文配图:OZ-TAL: Online Zero-Shot Temporal Action Localization
图 1 · 摘自论文原文
  • 利用现成视觉语言模型,不需训练直接推理
  • 在THUMOS14和ActivityNet上实现显著零样本性能提升
  • 适合实时视频分析中应对未知动作的场景

在线时序动作定位(On-TAL)旨在视频流持续输入时即时检测动作发生时间与类别。现有方法多依赖特定领域训练,面对未见过的动作泛化能力有限。本文提出在线零样本时序动作定位(OZ-TAL)新任务,设计一种无需训练的框架,基于现成的视觉语言模型(VLMs),并引入机制增强视觉表征、缓解其固有偏见。在THUMOS14和ActivityNet-1.3上建立新基准与基线,实验表明该方法在离线与在线零样本设置下均显著优于现有最先进方法。

原文摘要 · Abstract (English)

Online Temporal Action Localization (On-TAL) aims to detect the occurrence time and category of actions in untrimmed streaming videos immediately upon their completion. Recent advancements in this field focus on developing more sophisticated frameworks, shifting from Online Action Detection (OAD)-based aggregation paradigm to instance-level understanding. However, existing approaches are typically trained on specific domains and often exhibit limited generalization capabilities when applied to arbitrary videos, particularly in the presence of previously unseen actions. In this paper, we introduce a new task called Online Zero-shot Temporal Action Localization (OZ-TAL), which aims to detect previously unseen actions in an online fashion. Furthermore, we propose a training-free framework that leverages off-the-shelf Vision-Language Models (VLMs) while introducing additional mechanisms to enhance visual representations and mitigate their inherent biases. We establish new benchmarks and representative baselines for OZ-TAL on THUMOS14 and ActivityNet-1.3, and extensive experiments demonstrate that our method substantially outperforms existing state-of-the-art approaches under both offline and online zero-shot settings.

动作定位零样本在线推理视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。