通过频域增强语义建模,提升零样本骨骼动作识别准确率
Frequency-Semantic Enhanced Variational Autoencoder for Zero-Shot Skeleton-based Action Recognition

- 引入频域分解模块,区分高低频动作特征以增强语义学习
- 多层级对齐机制捕捉局部细节与全局对应,缩小语义鸿沟
- 校准交叉对齐损失,有效应对骨架与文本特征的模糊匹配问题
零样本骨骼动作识别旨在让模型识别训练中未见过的动作类别。现有方法多关注视觉与语义表示对齐,却忽视了语义空间中细微动作模式(如喝水与刷牙的手部动作差异)的重要性。为此,本文提出频率-语义增强变分自编码器(FS-VAE),通过频域分解探索骨骼语义表示学习。FS-VAE包含三个关键组件:1)基于频率的增强模块,通过高低频调整丰富骨骼语义学习,提升零样本识别鲁棒性;2)多层次语义动作描述机制,有效捕捉局部细节与全局对应关系,弥补骨骼序列中的信息损失;3)校准交叉对齐损失,使有效骨架-文本配对可抵消模糊配对的影响,缓解特征歧义。在多个基准数据集上的评估表明,频域增强的语义特征能有效区分视觉与语义相似的动作簇,显著提升零样本动作识别性能。
原文摘要 · Abstract (English)
Zero-shot skeleton-based action recognition aims to develop models capable of identifying actions beyond the categories encountered during training. Previous approaches have primarily focused on aligning visual and semantic representations but often overlooked the importance of fine-grained action patterns in the semantic space (e.g., the hand movements in drinking water and brushing teeth). To address these limitations, we propose a Frequency-Semantic Enhanced Variational Autoencoder (FS-VAE) to explore the skeleton semantic representation learning with frequency decomposition. FS-VAE consists of three key components: 1) a frequency-based enhancement module with high- and low-frequency adjustments to enrich the skeletal semantics learning and improve the robustness of zero-shot action recognition; 2) a semantic-based action description with multilevel alignment to capture both local details and global correspondence, effectively bridging the semantic gap and compensating for the inherent loss of information in skeleton sequences; 3) a calibrated cross-alignment loss that enables valid skeleton-text pairs to counterbalance ambiguous ones, mitigating discrepancies and ambiguities in skeleton and text features, thereby ensuring robust alignment. Evaluations on the benchmarks demonstrate the effectiveness of our approach, validating that frequency-enhanced semantic features enable robust differentiation of visually and semantically similar action clusters, improving zero-shot action recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。