arXiv:2512.18750cs.CV2025-12被引 1

通过多尺度时空注意力捕捉动作细节与整体流,提升视频动作识别精度

Context-Aware Network Based on Multi-scale Spatio-temporal Attention for Action Recognition in Videos

  • 设计双模块网络:多尺度时序模块提取不同速度运动,分组空间模块处理多尺度视觉特征
  • 在5个数据集上达到领先性能,最高准确率达88.4%(Diving48)
  • 适合需要精细动作理解的场景,如体育分析、医疗行为识别

动作识别是视频理解的关键任务,需全面捕捉跨尺度的时空线索。现有方法常忽略动作的多粒度特性。为此,本文提出上下文感知网络(CAN),包含两个核心模块:多尺度时序线索模块(MTCM)和分组空间线索模块(GSCM)。MTCM有效提取多尺度时序信息,捕获快速变化的运动细节与整体动作流;GSCM通过分组特征图并分别应用专用提取方法,实现多尺度空间特征建模。在五个基准数据集(Something-Something V1/V2、Diving48、Kinetics-400、UCF101)上的实验表明,该方法表现优异,优于多数主流模型:在Something-Something V1上达50.4%,V2上达63.9%,Diving48上达88.4%,Kinetics-400上达74.9%,UCF101上达86.9%。结果凸显了多尺度时空线索对鲁棒动作识别的重要性。

原文摘要 · Abstract (English)

Action recognition is a critical task in video understanding, requiring the comprehensive capture of spatio-temporal cues across various scales. However, existing methods often overlook the multi-granularity nature of actions. To address this limitation, we introduce the Context-Aware Network (CAN). CAN consists of two core modules: the Multi-scale Temporal Cue Module (MTCM) and the Group Spatial Cue Module (GSCM). MTCM effectively extracts temporal cues at multiple scales, capturing both fast-changing motion details and overall action flow. GSCM, on the other hand, extracts spatial cues at different scales by grouping feature maps and applying specialized extraction methods to each group. Experiments conducted on five benchmark datasets (Something-Something V1 and V2, Diving48, Kinetics-400, and UCF101) demonstrate the effectiveness of CAN. Our approach achieves competitive performance, outperforming most mainstream methods, with accuracies of 50.4% on Something-Something V1, 63.9% on Something-Something V2, 88.4% on Diving48, 74.9% on Kinetics-400, and 86.9% on UCF101. These results highlight the importance of capturing multi-scale spatio-temporal cues for robust action recognition.

动作识别时空建模多尺度视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。