通过多层次关系建模,提升少样本动作识别的泛化能力
Hierarchical Relation-augmented Representation Generalization for Few-shot Action Recognition
- 构建帧间、视频间、任务间三层关系建模框架
- 在5个数据集上显著超越现有最优方法
- 适合需要跨任务复用时序知识的研究者
少样本动作识别(FSAR)旨在仅用少量样本识别新动作类别。现有方法通常通过设计帧间时间建模或粗粒度视频级交互来学习每段视频的帧级表征,但它们将每个任务独立处理,忽视了视频间的细粒度时间关系,无法捕捉跨视频共享的时间模式,也无法复用历史任务中的时序知识。为此,我们提出HR2G-shot框架,统一建模三种关系:帧间、视频间和任务间,从整体视角学习任务特异性时间模式。除帧间交互外,还设计两个组件:一是视频间语义相关性(ISC),以细粒度方式实现跨视频帧级交互,增强类内一致性与类间可分性;二是任务间知识迁移(IKT),从存储历史任务多样时序模式的记忆库中检索并聚合相关知识。在五个基准上的大量实验表明,HR2G-shot优于当前领先的FSAR方法。
原文摘要 · Abstract (English)
Few-shot action recognition (FSAR) aims to recognize novel action categories with few exemplars. Existing methods typically learn frame-level representations for each video by designing inter-frame temporal modeling strategies or inter-video interaction at the coarse video-level granularity. However, they treat each episode task in isolation and neglect fine-grained temporal relation modeling between videos, thus failing to capture shared fine-grained temporal patterns across videos and reuse temporal knowledge from historical tasks. In light of this, we propose HR2G-shot, a Hierarchical Relation-augmented Representation Generalization framework for FSAR, which unifies three types of relation modeling (inter-frame, inter-video, and inter-task) to learn task-specific temporal patterns from a holistic view. Going beyond conducting inter-frame temporal interactions, we further devise two components to respectively explore inter-video and inter-task relationships: i) Inter-video Semantic Correlation (ISC) performs cross-video frame-level interactions in a fine-grained manner, thereby capturing task-specific query features and enhancing both intra-class consistency and inter-class separability; ii) Inter-task Knowledge Transfer (IKT) retrieves and aggregates relevant temporal knowledge from the bank, which stores diverse temporal patterns from historical episode tasks. Extensive experiments on five benchmarks show that HR2G-shot outperforms current top-leading FSAR methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。