提出自适应对齐多尺度动作特征,解决视频时序错位难题
A$^2$M$^2$-Net: Adaptively Aligned Multi-Scale Moment for Few-Shot Action Recognition
- 用多尺度二阶统计量捕捉视频动态特征
- 自适应选择关键特征,提升对时序错位的鲁棒性
- 在5个数据集上超越主流方法,适合少样本动作识别任务
为降低大规模标注成本,少样本动作识别(FSAR)近年受到广泛关注。现有方法常忽视个体运动模式差异,且未充分挖掘视频动态的特征统计信息,导致难以应对2D骨干网络下的严重时序错位问题。为此,本文提出自适应对齐多尺度二阶矩网络(A²M²-Net),通过构建一组强表示候选并以实例引导方式自适应对齐,有效描述潜在视频动态。核心包含两个组件:自适应对齐(A²模块)用于匹配,多尺度二阶矩(M²块)用于生成多时空尺度的语义二阶描述符。其中,M²块在多个时空尺度上生成语义二阶描述符;A²模块则根据个体运动模式自适应选择重要候选。该方法建立自适应对齐机制,显著提升对时序错位的处理能力。实验在5个常用FSAR基准上进行,结果表明,A²M²-Net性能媲美甚至超越当前最先进方法,展现出优异的有效性与泛化能力。
原文摘要 · Abstract (English)
Thanks to capability to alleviate the cost of large-scale annotation, few-shot action recognition (FSAR) has attracted increased attention of researchers in recent years. Existing FSAR approaches typically neglect the role of individual motion pattern in comparison, and under-explore the feature statistics for video dynamics. Thereby, they struggle to handle the challenging temporal misalignment in video dynamics, particularly by using 2D backbones. To overcome these limitations, this work proposes an adaptively aligned multi-scale second-order moment network, namely A$^2$M$^2$-Net, to describe the latent video dynamics with a collection of powerful representation candidates and adaptively align them in an instance-guided manner. To this end, our A$^2$M$^2$-Net involves two core components, namely, adaptive alignment (A$^2$ module) for matching, and multi-scale second-order moment (M$^2$ block) for strong representation. Specifically, M$^2$ block develops a collection of semantic second-order descriptors at multiple spatio-temporal scales. Furthermore, A$^2$ module aims to adaptively select informative candidate descriptors while considering the individual motion pattern. By such means, our A$^2$M$^2$-Net is able to handle the challenging temporal misalignment problem by establishing an adaptive alignment protocol for strong representation. Notably, our proposed method generalizes well to various few-shot settings and diverse metrics. The experiments are conducted on five widely used FSAR benchmarks, and the results show our A$^2$M$^2$-Net achieves very competitive performance compared to state-of-the-arts, demonstrating its effectiveness and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。