用视觉语言模型实现零样本开放词汇动作分割,无需训练即可识别新动作。
Exploring Vision-Language Models for Open-Vocabulary Zero-Shot Action Segmentation
- 基于帧-动作嵌入相似性匹配候选动作标签
- 在标准数据集上零样本性能媲美有监督方法
- 首次系统评估14种VLM在开放词汇动作分割中的表现
时间动作分割(TAS)需将视频划分为动作片段,但活动种类繁多且划分方式多样,难以构建全覆盖数据集。现有方法受限于封闭词汇和固定标签集。本文探索尚未充分研究的开放词汇零样本时间动作分割(OVTAS)问题,利用视觉语言模型(VLMs)的强零样本能力。提出无训练流水线:帧-动作嵌入相似性(FAES)将视频帧与候选动作标签匹配,相似度矩阵时间分割(SMTS)保证时序一致性。除提出OVTAS外,系统评估14种不同VLM,首次提供其在开放词汇动作分割中的适用性分析。在标准基准测试中,OVTAS在无任务特定监督下取得优异结果,验证了VLM在结构化时序理解中的潜力。
原文摘要 · Abstract (English)
Temporal Action Segmentation (TAS) requires dividing videos into action segments, yet the vast space of activities and alternative breakdowns makes collecting comprehensive datasets infeasible. Existing methods remain limited to closed vocabularies and fixed label sets. In this work, we explore the largely unexplored problem of Open-Vocabulary Zero-Shot Temporal Action Segmentation (OVTAS) by leveraging the strong zero-shot capabilities of Vision-Language Models (VLMs). We introduce a training-free pipeline that follows a segmentation-by-classification design: Frame-Action Embedding Similarity (FAES) matches video frames to candidate action labels, and Similarity-Matrix Temporal Segmentation (SMTS) enforces temporal consistency. Beyond proposing OVTAS, we present a systematic study across 14 diverse VLMs, providing the first broad analysis of their suitability for open-vocabulary action segmentation. Experiments on standard benchmarks show that OVTAS achieves strong results without task-specific supervision, underscoring the potential of VLMs for structured temporal understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。