让文字更聚焦动作区域,提升未见动作的精准定位
TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detection

- 用动作集中度聚合视频片段,生成前景加权嵌入
- 在多个数据集上超越现有方法,跨数据集泛化能力更强
- 适合需要高精度零样本动作检测的视频分析场景
零样本时间动作检测(ZSTAD)旨在定位并识别未见动作类别在非剪辑视频中的实例。尽管现有方法通过改进文本-视频对齐架构已取得一定效果,但仍难以捕捉动作类别间的语义差异,导致产生无关文本预测。为此,我们提出一种面向零样本时间动作检测的文本-前景聚焦对齐方法(TF-CADE),显式将文本信息与动作相关的前景区域对齐。具体地,引入动作集中度聚合(ACA),通过提取动作集中度得分,将时序信息丰富的视频片段聚合为前景加权的视频嵌入。该前景聚焦对齐增强了文本与视频特征间的语义一致性,并提升了类间可区分性。此外,基于确定性的置信度重加权(CCR)策略,利用前景感知相似性优化每个片段的置信度分数,在推理阶段有效抑制无关动作类别。大量实验表明,所提方法不仅在分布内设置下达到当前最优性能,还在跨数据集泛化至未见动作类别方面表现优异。
原文摘要 · Abstract (English)
Zero-Shot Temporal Action Detection (ZSTAD) aims to lo- calize and recognize action instances from unseen action categories in untrimmed videos. Although existing meth- ods have shown effectiveness by advancing architectural text-video alignment, they still struggle with capturing se- mantic distinctions between action classes, resulting in text- irrelevant predictions. To address this issue, we propose a Text-Foreground Concentrated Alignment for zero-shot temporal action DEtector (TF-CADE) that explicitly aligns textual information with action-relevant foreground regions. Specifically, we introduce Action Concentrate Aggregation (ACA), which extracts action concentrate scores to aggregate temporally informative video segments into a foreground- weighted video embedding. This foreground concentrated alignment enhances the semantic consistency between text and video features and improves inter-class discriminabil- ity. In addition, a Certainty-based Confidence Re-weighting (CCR) strategy refines per-snippet confidence scores by lever- aging foreground-aware similarity, effectively suppressing irrelevant action classes during inference. Extensive evalua- tions show that our TF-CADE not only achieves state-of-the- art performance under in-distribution settings but also excels in cross-dataset generalization to unseen action classes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。