通过恒等约束提升骨骼动作识别的无监督表示学习效果
Idempotent Unsupervised Representation Learning for Skeleton-Based Action Recognition

- 引入恒等性约束增强特征空间一致性,保留运动语义关键信息
- NTU RGB+D数据集上准确率从84.6%提升至86.2%
- 在零样本场景下也能识别原模型无法处理的动作
生成模型在图像生成中表现优异,逐渐被用于识别任务。然而,在骨骼动作识别中,现有预训练生成方法提取的特征包含与识别无关的冗余信息,违背了骨骼数据空间稀疏、时间连续的特性,导致性能下降。为此,我们提出一种新型骨骼基恒等生成模型(IGM),用于无监督表示学习。理论上证明生成模型与最大熵编码等价,为通过对比学习压缩特征提供理论支持。进一步引入恒等性约束,在特征空间施加强一致性正则化,使特征仅保留运动语义的关键信息。在基准数据集NTU RGB+D和PKUMMD上的大量实验表明,该方法有效:在NTU 60 xsub上,准确率从84.6%提升至86.2%;在零样本迁移场景中,对先前无法识别的动作也取得显著成效。
原文摘要 · Abstract (English)
Generative models, as a powerful technique for generation, also gradually become a critical tool for recognition tasks. However, in skeleton-based action recognition, the features obtained from existing pre-trained generative methods contain redundant information unrelated to recognition, which contradicts the nature of the skeleton's spatially sparse and temporally consistent properties, leading to undesirable performance. To address this challenge, we make efforts to bridge the gap in theory and methodology and propose a novel skeleton-based idempotent generative model (IGM) for unsupervised representation learning. More specifically, we first theoretically demonstrate the equivalence between generative models and maximum entropy coding, which demonstrates a potential route that makes the features of generative models more compact by introducing contrastive learning. To this end, we introduce the idempotency constraint to form a stronger consistency regularization in the feature space, to push the features only to maintain the critical information of motion semantics for the recognition task. Our extensive experiments on benchmark datasets, NTU RGB+D and PKUMMD, demonstrate the effectiveness of our proposed method. On the NTU 60 xsub dataset, we observe a performance improvement from 84.6$\%$ to 86.2$\%$. Furthermore, in zero-shot adaptation scenarios, our model demonstrates significant efficacy by achieving promising results in cases that were previously unrecognizable. Our project is available at \url{https://github.com/LanglandsLin/IGM}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。