arXiv:2603.24030cs.CVcs.MM2026-03中稿 · CVPR

通过分阶段分解与对齐,提升未见动作的时序检测效果

Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection

  • 用大模型思维链拆解动作标签为阶段描述
  • 自适应筛选各阶段相关视频片段,对齐语义与视觉
  • 分阶段对齐增强知识迁移,适合开放词汇检测研究者

开放词汇时序动作检测(OV-TAD)旨在识别未修剪视频中未见过的动作类别。现有方法仅依赖标签语义与视觉特征的全局对齐,难以有效传递已见类别的时序一致视觉知识。为此,本文提出分阶段分解与对齐(PDA)框架,实现细粒度动作模式学习以促进先验知识迁移。首先引入基于思维链(CoT)的语义分解模块(CSD),利用大语言模型能力自动将动作标签分解为连贯的阶段级描述,模拟人类认知过程;随后设计文本引导前景过滤(TIF)模块,基于阶段语义线索自适应筛选每个阶段的动作相关片段,生成语义对齐的视觉表征;进一步提出自适应分阶段对齐(APA)模块,实现阶段级视觉-文本匹配,并自适应聚合各阶段对齐结果进行最终预测。该机制有助于捕捉可迁移的动作模式,显著提升对未见动作的泛化能力。在两个OV-TAD基准数据集上的大量实验验证了所提方法的优越性。

原文摘要 · Abstract (English)

Open-Vocabulary Temporal Action Detection (OV-TAD) aims to classify and localize action segments in untrimmed videos for unseen categories. Previous methods rely solely on global alignment between label-level semantics and visual features, which is insufficient to transfer temporal consistent visual knowledge from seen to unseen classes. To address this, we propose a Phase-wise Decomposition and Alignment (PDA) framework, which enables fine-grained action pattern learning for effective prior knowledge transfer. Specifically, we first introduce the CoT-Prompting Semantic Decomposition (CSD) module, which leverages the chain-of-thought (CoT) reasoning ability of large language models to automatically decompose action labels into coherent phase-level descriptions, emulating human cognitive processes. Then, Text-infused Foreground Filtering (TIF) module is introduced to adaptively filter action-relevant segments for each phase leveraging phase-wise semantic cues, producing semantically aligned visual representations. Furthermore, we propose the Adaptive Phase-wise Alignment (APA) module to perform phase-level visual-textual matching, and adaptively aggregates alignment results across phases for final prediction. This adaptive phase-wise alignment facilitates the capture of transferable action patterns and significantly enhances generalization to unseen actions. Extensive experiments on two OV-TAD benchmarks demonstrated the superiority of the proposed method.

动作检测开放词汇提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。