arXiv:2601.07666cs.CVcs.AI2026-01

用概率建模提升骨架动作识别的自监督学习效果

Variational Contrastive Learning for Skeleton-based Action Recognition

  • 将变分推断融入对比学习,建模人体动作的不确定性
  • 在低标注数据下仍优于现有方法,跨数据集泛化性强
  • 生成更关注关键关节的语义特征,适合动作识别任务

近年来,基于骨架的动作识别自监督表示学习随着对比学习方法的发展取得进展。然而,多数对比范式本质上是判别式的,难以捕捉人体运动固有的可变性和不确定性。为此,我们提出一种变分对比学习框架,将概率潜在建模与对比自监督学习结合。该框架能学习到结构化且语义清晰的表示,在三个广泛使用的骨架动作识别基准上均表现优异,尤其在低标签场景下优势明显。定性分析表明,相比其他方法,本方法生成的特征对运动和样本特性更敏感,注意力更集中于重要骨骼关节点。

原文摘要 · Abstract (English)

In recent years, self-supervised representation learning for skeleton-based action recognition has advanced with the development of contrastive learning methods. However, most of contrastive paradigms are inherently discriminative and often struggle to capture the variability and uncertainty intrinsic to human motion. To address this issue, we propose a variational contrastive learning framework that integrates probabilistic latent modeling with contrastive self-supervised learning. This formulation enables the learning of structured and semantically meaningful representations that generalize across different datasets and supervision levels. Extensive experiments on three widely used skeleton-based action recognition benchmarks show that our proposed method consistently outperforms existing approaches, particularly in low-label regimes. Moreover, qualitative analyses show that the features provided by our method are more relevant given the motion and sample characteristics, with more focus on important skeleton joints, when compared to the other methods.

动作识别对比学习变分模型自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。