arXiv:2502.12582cs.CV2025-02被引 3

提出自适应原型模型,提升多属性少样本动作识别准确率

Adaptive Prototype Model for Attribute-based Multi-label Few-shot Action Recognition

  • 用文本约束模块构建属性原型,融合语义信息
  • 在多属性少样本任务上达到当前最优性能
  • 适合需要细粒度动作理解的智能视频分析场景

现实世界中的动作识别系统需整合多种属性以全面理解人类行为,但单一模型同时识别多个属性易导致精度下降。本文提出自适应属性原型模型(AAPM),通过引入文本约束模块(TCM)融合潜在标签的文本信息,约束不同属性原型表示的构建,并设计属性分配方法(AAM)缓解训练偏差,提升鲁棒性。此外,构建了新的基于属性的多标签视频数据集Multi-Kinetics,包含动作、场景、物体等多样化属性标签。大量实验表明,AAPM在多属性少样本动作识别和单属性少样本动作识别任务上均达到领先水平。

原文摘要 · Abstract (English)

In real-world action recognition systems, incorporating more attributes helps achieve a more comprehensive understanding of human behavior. However, using a single model to simultaneously recognize multiple attributes can lead to a decrease in accuracy. In this work, we propose a novel method i.e. Adaptive Attribute Prototype Model (AAPM) for human action recognition, which captures rich action-relevant attribute information and strikes a balance between accuracy and robustness. Firstly, we introduce the Text-Constrain Module (TCM) to incorporate textual information from potential labels, and constrain the construction of different attributes prototype representations. In addition, we explore the Attribute Assignment Method (AAM) to address the issue of training bias and increase robustness during the training process.Furthermore, we construct a new video dataset with attribute-based multi-label called Multi-Kinetics for evaluation, which contains various attribute labels (e.g. action, scene, object, etc.) related to human behavior. Extensive experiments demonstrate that our AAPM achieves the state-of-the-art performance in both attribute-based multi-label few-shot action recognition and single-label few-shot action recognition. The project and dataset are available at an anonymous account https://github.com/theAAPM/AAPM

少样本学习动作识别多标签原型网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。