arXiv:2503.14430cs.CV2025-03中稿 · Computer Vision an…被引 3

聚焦动作实例,提升少样本动作识别准确率

Joint Image-Instance Spatial-Temporal Attention for Few-shot Action Recognition

  • 通过文本引导分割感知动作相关实例
  • 融合图像与实例特征构建时空注意力关系
  • 适用于小样本场景下动作识别任务

少样本动作识别(FSAR)是计算机视觉中的关键挑战,要求从极少量样例中识别动作。现有方法多依赖图像级特征构建时序依赖并生成类别原型,但这些特征常包含背景噪声,且对真实前景(动作相关实例)关注不足,尤其在少样本情况下限制了识别能力。为此,我们提出一种新的联合图像-实例级时空注意力方法(I2ST)。其核心思想是感知动作相关实例,并通过时空注意力将它们与图像特征融合。I2ST包含两个关键组件:动作相关实例感知和联合图像-实例时空注意力。给定特征提取器的基础表示后,动作相关实例感知在文本引导分割模型的指导下识别动作相关实例;随后,联合图像-实例时空注意力用于构建实例与图像间的特征依赖关系。

原文摘要 · Abstract (English)

Few-shot Action Recognition (FSAR) constitutes a crucial challenge in computer vision, entailing the recognition of actions from a limited set of examples. Recent approaches mainly focus on employing image-level features to construct temporal dependencies and generate prototypes for each action category. However, a considerable number of these methods utilize mainly image-level features that incorporate background noise and focus insufficiently on real foreground (action-related instances), thereby compromising the recognition capability, particularly in the few-shot scenario. To tackle this issue, we propose a novel joint Image-Instance level Spatial-temporal attention approach (I2ST) for Few-shot Action Recognition. The core concept of I2ST is to perceive the action-related instances and integrate them with image features via spatial-temporal attention. Specifically, I2ST consists of two key components: Action-related Instance Perception and Joint Image-Instance Spatial-temporal Attention. Given the basic representations from the feature extractor, the Action-related Instance Perception is introduced to perceive action-related instances under the guidance of a text-guided segmentation model. Subsequently, the Joint Image-Instance Spatial-temporal Attention is used to construct the feature dependency between instances and images...

少样本识别动作识别时空注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。