用生成模型先验提升零样本动作识别,让文字与骨骼动作对齐更准。
GenPrior: Unleashing Text-to-Motion Generative Priors for Zero-Shot Skeleton-based Action Recognition

- 用文本到动作生成模型提取带结构的动作先验
- 在NTU-60/120和PKU-MMD上超越现有方法
- 适合做零样本动作识别或生成模型迁移的研究者
零样本骨骼动作识别(ZSAR)旨在通过将骨骼特征与文本语义对齐来识别未见过的动作类别。然而,现有方法依赖于缺乏几何结构和物理约束的文本导出原型,导致显著的语义-运动间隙。为此,我们提出首个利用预训练文本到动作(T2M)模型生成先验的框架GenPrior。具体地,我们引入分散门控特征融合,从生成的运动序列中提炼运动原型和类内差异,并通过学习的门控网络自适应地将可靠结构信息注入文本嵌入,同时抑制合成伪影。此外,我们提出生成原型精炼机制,利用这些增强的原型作为锚点挖掘高置信度的未见样本,校准类别原型以逼近真实分布,从而释放性能优势。在NTU-60、NTU-120和PKU-MMD上的大量实验表明,GenPrior在零样本和广义零样本设置下均达到最优性能。代码已开源。
原文摘要 · Abstract (English)
Zero-shot skeleton-based action recognition (ZSAR) aims to recognize unseen action categories by aligning skeleton features with textual semantics. However, existing methods rely on text-derived prototypes that inherently lack geometric structure and physical constraints, resulting in a pronounced \textit{semantic-kinematic gap}. To bridge this gap, we propose \textbf{GenPrior}, the first framework to exploit generative priors from pre-trained Text-to-Motion (T2M) models for ZSAR. Specifically, we introduce Dispersion-Gated Feature Fusion, which distills kinematic prototypes and intra-class dispersion from generative motion sequences and employs a learned gating network to adaptively inject reliable structural cues into textual embeddings while suppressing synthetic artifacts. Furthermore, we propose Generative Prototype Refinement, which leverages these generation-enhanced prototypes as anchors to mine high-confidence unseen samples, calibrating class prototypes toward the true distribution and thereby unleashing strong performance gains. Extensive experiments on NTU-60, NTU-120, and PKU-MMD demonstrate that GenPrior achieves state-of-the-art performance under both zero-shot and generalized zero-shot settings. Code is available at https://github.com/jidongkuang/GenPrior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。