arXiv:2604.09164cs.CV2026-04被引 1

用高效时空聚焦模块提升长视频动作检测精度

Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection

  • 设计轻量级时空聚焦适配器,融合边界感知状态空间模型
  • 在多个数据集上实现更高定位准确率,显著优于传统方法
  • 适合需要高效长视频分析的工业场景与实时系统

时序人体动作检测旨在识别和定位未剪辑视频中的动作片段,是视频理解的关键任务。尽管卷积网络与变换器架构已取得进展,但在处理长视频序列时仍面临特征冗余与全局依赖建模能力下降的问题,严重限制了其在真实场景中的可扩展性。状态空间模型(SSMs)具备线性长程建模与强全局时序推理能力,具有潜力。本文重新思考SSM在时序建模中的应用,提出一种新框架用于视频人体动作检测。具体地,我们在预训练层中引入高效时空聚焦(ESTF)适配器,该模块结合了所提出的时序边界感知状态空间模型(TB-SSM)进行时序特征建模,并高效处理空间特征。我们在多个基准上进行了全面定量分析,对比了本方法与以往基于SSM及其他结构的方法。大量实验表明,所提策略显著提升了定位性能与鲁棒性,验证了方法的有效性。

原文摘要 · Abstract (English)

Temporal human action detection aims to identify and localize action segments within untrimmed videos, serving as a pivotal task in video understanding. Despite the progress achieved by prior architectures like CNN and Transformer models, these continue to struggle with feature redundancy and degraded global dependency modeling capabilities when applied to long video sequences. These limitations severely constrain their scalability in real-world video analysis. State Space Models (SSMs) offer a promising alternative with linear long-term modeling and robust global temporal reasoning capabilities. Rethinking the application of SSMs in temporal modeling, this research constructs a novel framework for video human action detection. Specifically, we introduce the Efficient Spatial-Temporal Focal (ESTF) Adapter into the pre-trained layers. This module integrates the advantages of our proposed Temporal Boundary-aware SSM(TB-SSM) for temporal feature modeling with efficient processing of spatial features. We perform comprehensive and quantitative analyses across multiple benchmarks, comparing our proposed method against previous SSM-based and other structural methods. Extensive experiments demonstrate that our improved strategy significantly enhances both localization performance and robustness, validating the effectiveness of our proposed method.

动作检测状态空间模型视频理解轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。