用局部全局联合建模提升零样本动作定位精度
ConTrans: Learning Text-enhanced Local-global Temporal Representations for Zero-shot Temporal Action Localization

- 融合卷积与自注意力,捕捉帧间局部相关性与长程上下文
- 在ActivityNet-1.3和THUMOS14上达到新最优性能
- 适合关注视频理解中零样本检测的研究者
零样本时间动作定位(ZS-TAL)旨在检测未见过的动作。现有方法多聚焦长程上下文建模,忽视了帧间相对偏移带来的局部相关性。同时,其性能受限于浅层网络结构导致的特征表达能力不足。本文提出新型多尺度编码器ConTrans,将卷积归纳偏置与Transformer自注意力结合,协同捕捉细粒度局部依赖与长程全局上下文,实现更全面的特征表示。在ActivityNet-1.3和THUMOS14数据集上的实验表明,ConTrans显著优于现有方法,建立了新的基准。
原文摘要 · Abstract (English)
Zero-shot Temporal Action Localization (ZS-TAL) aims to detect and locate previously unseen actions in untrimmed videos. However, existing approaches primarily focus on modeling long-range contextual information, often neglecting the critical relative-offset-based local correlations between video frames. Furthermore, their performance is hindered by limited feature representation capabilities due to the shallow nature of their network architectures. In this paper, we address these limitations by introducing a novel local-global multi-scale feature representation module. We propose a novel multi-scale encoder architecture, termed ConTrans, that integrates convolutional (Conv) inductive biases with transformer Self-attention to jointly capture fine-grained local dependencies and long-range global context, leading to more comprehensive feature representations than existing methods. Experimental evaluations on the ActivityNet-1.3 and THUMOS14 datasets demonstrate that ConTrans significantly outperforms existing methods, establishing a new benchmark for ZS-TAL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。