arXiv:2511.10091cs.CV2025-11中稿 · AAAI被引 5

用视觉运动知识指导骨骼表征学习,提升动作识别效果。

SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition

  • 通过视频模型生成视觉运动知识,监督骨骼表征学习
  • 在多个基准上实现优于线性方法的分类性能
  • 支持零样本场景,适合动作理解与描述任务

大型语言模型(LLM)具备丰富的隐式知识和强大的迁移能力。本文探索将LLM与人体骨骼结合用于动作分类与描述。然而,当将LLM作为识别器时,面临两个问题:1)如何让LLM理解骨骼数据?2)如何区分不同动作?为此,我们提出一种新范式SUGAR(基于视觉-运动知识学习骨骼表征的动作识别)。该方法首先利用现成的大规模视频模型作为知识库,生成与动作相关的视觉和运动信息;随后,通过这些先验知识监督骨骼表征学习,获得离散表示;最后,使用未经微调的LLM理解这些表示并生成动作目标与描述。特别地,我们设计了时间查询投影(TQP)模块,持续建模长序列骨骼信号。在多个基于骨骼的动作分类基准上验证了SUGAR的有效性。此外,零样本实验表明,SUGAR比线性方法更具泛化能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) hold rich implicit knowledge and powerful transferability. In this paper, we explore the combination of LLMs with the human skeleton to perform action classification and description. However, when treating LLM as a recognizer, two questions arise: 1) How can LLMs understand skeleton? 2) How can LLMs distinguish among actions? To address these problems, we introduce a novel paradigm named learning Skeleton representation with visUal-motion knowledGe for Action Recognition (SUGAR). In our pipeline, we first utilize off-the-shelf large-scale video models as a knowledge base to generate visual, motion information related to actions. Then, we propose to supervise skeleton learning through this prior knowledge to yield discrete representations. Finally, we use the LLM with untouched pre-training weights to understand these representations and generate the desired action targets and descriptions. Notably, we present a Temporal Query Projection (TQP) module to continuously model the skeleton signals with long sequences. Experiments on several skeleton-based action classification benchmarks demonstrate the efficacy of our SUGAR. Moreover, experiments on zero-shot scenarios show that SUGAR is more versatile than linear-based methods.

动作识别骨骼表征大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。