通过分层度量学习,提升少样本视频动作识别的泛化能力
Few-Shot Video Recognition via Hierarchical Metric Learning

- 构建时空特征处理流程,增强跨帧全局空间表征
- 多阶段度量约束联合优化,显著提升类别原型鲁棒性
- 在5个主流数据集上验证,对小样本场景效果突出
少样本动作识别(FSAR)旨在仅用少量标注视频样本识别未见动作类别。现有方法通常在网络输出层采用单原型监督,难以充分挖掘视频中丰富的跨帧全局空间信息。即使已有多层次度量方案,也仅在中间层并行施加原型约束,缺乏沿完整特征路径的渐进式监督,导致学习到的类别原型泛化能力有限。为此,我们提出一种新方法——分层度量学习用于少样本动作识别(HML-FSAR)。首先,设计空间增强模块以捕捉跨帧全局空间表示,结合时序多头注意力、异构对齐、时空特征融合与词典学习模块,构建完整特征处理流程。其次,将分层度量学习(HML)嵌入至HML-FSAR中,包含中心度量、对齐度量、对比度量、词典度量和原型度量,从帧级表示到最终类别原型逐级施加互补约束,联合优化特征紧凑性、异构时空对齐性、类间可区分性及抗噪鲁棒性。所提方法在五个常用FSAR数据集上验证,实验结果充分证明其有效性。
原文摘要 · Abstract (English)
Few-shot action recognition (FSAR) aims to recognize unseen action categories with only a small number of annotated video samples. Recent works typically apply single-prototype supervision at the network output and fail to sufficiently exploit rich cross-frame global spatial information in videos. Even existing multi-level metric schemes only impose parallel prototype constraints on intermediate layers, without progressive supervision along the full feature pipeline, which results in limited generalization ability of the learned class prototypes. Inspired by this, we present a novel method, hierarchical metric learning for few-shot action recognition (HML-FSAR). First, a spatial-enhanced module is developed to capture cross-frame global spatial representations. Combined with temporal MHA, heterogeneous alignment, spatial-temporal feature fusion and dictionary learning modules, it constructs the complete feature processing pipeline. Second, a hierarchical metric learning (HML) strategy is embedded into HML-FSAR. Composed of center metric, alignment metric, contrastive metric, dictionary metric and prototype metric, HML imposes progressive multi-stage complementary constraints from frame-level representations to final class prototypes, so as to jointly optimize feature compactness, heterogeneous spatial-temporal alignment, inter-class discriminability and anti-noise robustness. The proposed HML-FSAR method is validated on five widely-used FSAR datasets, and experimental results fully demonstrate its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。