提出新框架,让模型零样本识别动作更准更快。
SCALE: Semantic- and Confidence-Aware Conditional Variational Autoencoder for Zero-shot Skeleton-based Action Recognition
- 用能量排序替代对齐,直接评估未知动作类别的似然
- 在NTU-60和NTU-120上优于主流VAE与对齐方法
- 适合做零样本动作识别且无需生成样本,推理快
零样本骨架动作识别(ZSAR)旨在不依赖目标类别的训练骨架数据,仅通过文本辅助语义进行动作识别。现有方法常依赖显式的骨架-文本对齐,但在动作名称模糊或未见类别语义相近时表现脆弱。本文提出SCALE,一种轻量级、确定性的基于列表能量的语义与置信度感知框架,将ZSAR建模为类别条件能量排序问题。该框架构建了一个文本条件的变分自编码器,其中冻结的文本表征同时参数化潜空间先验与解码器,可在测试时不生成样本的情况下实现基于似然的未见类别评估。为区分竞争假设,引入语义与置信度感知的列表能量损失,强调语义相似的难负例,并利用后验不确定性自适应调整决策边界,重加权模糊训练实例。此外,采用潜原型对比目标,使后验均值与文本导出的潜原型对齐,提升语义组织与类别可分性,无需直接特征匹配。在NTU-60和NTU-120数据集上的实验表明,SCALE持续优于先前基于VAE与对齐的方法,且与扩散模型方法相当。
原文摘要 · Abstract (English)
Zero-shot skeleton-based action recognition (ZSAR) aims to recognize action classes without any training skeletons from those classes, relying instead on auxiliary semantics from text. Existing approaches frequently depend on explicit skeleton-text alignment, which can be brittle when action names underspecify fine-grained dynamics and when unseen classes are semantically confusable. We propose SCALE, a lightweight and deterministic Semantic- and Confidence-Aware Listwise Energy-based framework that formulates ZSAR as class-conditional energy ranking. SCALE builds a text-conditioned Conditional Variational Autoencoder where frozen text representations parameterize both the latent prior and the decoder, enabling likelihood-based evaluation for unseen classes without generating samples at test time. To separate competing hypotheses, we introduce a semantic- and confidence-aware listwise energy loss that emphasizes semantically similar hard negatives and incorporates posterior uncertainty to adapt decision margins and reweight ambiguous training instances. Additionally, we utilize a latent prototype contrast objective to align posterior means with text-derived latent prototypes, improving semantic organization and class separability without direct feature matching. Experiments on NTU-60 and NTU-120 datasets show that SCALE consistently improves over prior VAE- and alignment-based baselines while remaining competitive with diffusion-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。