用语言上下文增强骨骼动作表示,提升零样本识别能力
SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognition
- 通过语言模型生成上下文提示,丰富骨骼运动特征
- 在多个基准上达到领先性能,尤其擅长区分相似动作
- 适合需要细粒度动作识别的零样本场景
零样本骨骼动作识别旨在通过语义描述将已见类别的知识迁移至未见动作。现有方法通常在共享隐空间中对齐骨骼特征与文本嵌入,但缺乏动作相关物体等上下文线索,导致骨骼与语义表征间存在固有鸿沟,难以区分视觉相似动作。为此,我们提出SkeletonContext,一种基于提示的框架,通过语言驱动的上下文语义增强骨骼运动表征。具体地,设计了跨模态上下文提示模块,利用预训练语言模型在LLM引导下重建被掩码的上下文提示,有效将语言上下文传递至骨骼编码器,实现实例级语义定位与跨模态对齐。同时引入关键关节解耦模块,分离与运动相关的关节特征,确保在无显式物体交互时仍具备鲁棒的动作理解能力。在多个基准上的大量实验表明,SkeletonContext在常规与广义零样本设置下均达到当前最优表现,验证了其在上下文推理与细粒度动作区分方面的有效性。
原文摘要 · Abstract (English)
Zero-shot skeleton-based action recognition aims to recognize unseen actions by transferring knowledge from seen categories through semantic descriptions. Most existing methods typically align skeleton features with textual embeddings within a shared latent space. However, the absence of contextual cues, such as objects involved in the action, introduces an inherent gap between skeleton and semantic representations, making it difficult to distinguish visually similar actions. To address this, we propose SkeletonContext, a prompt-based framework that enriches skeletal motion representations with language-driven contextual semantics. Specifically, we introduce a Cross-Modal Context Prompt Module, which leverages a pretrained language model to reconstruct masked contextual prompts under guidance derived from LLMs. This design effectively transfers linguistic context to the skeleton encoder for instance-level semantic grounding and improved cross-modal alignment. In addition, a Key-Part Decoupling Module is incorporated to decouple motion-relevant joint features, ensuring robust action understanding even in the absence of explicit object interactions. Extensive experiments on multiple benchmarks demonstrate that SkeletonContext achieves state-of-the-art performance under both conventional and generalized zero-shot settings, validating its effectiveness in reasoning about context and distinguishing fine-grained, visually similar actions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。