用多骨骼结构对比学习提升动作识别泛化能力
MS-CLR: Multi-Skeleton Contrastive Learning for Human Action Recognition
- 从同一视频中提取多种骨骼结构,通过对比学习对齐表示
- 在NTU RGB+D数据集上超越单骨骼基线,达新最好结果
- 适合需要跨数据集泛化的动作识别研究者
对比学习在基于骨架的动作识别中备受关注,因其能从无标签数据中学习鲁棒表征。然而,现有方法依赖单一骨骼结构,限制了其在不同关节布局和解剖覆盖范围数据集上的泛化能力。本文提出多骨骼对比学习(MS-CLR),一种通用自监督框架,通过在同一序列中提取的多种骨骼规范对齐姿态表征,促使模型学习结构不变性并捕捉多样化的解剖线索,从而获得更丰富且更具泛化性的特征。为此,我们改进了ST-GCN架构,通过统一表示方案处理具有不同关节布局和尺度的骨骼。在NTU RGB+D 60和120数据集上的实验表明,MS-CLR始终优于强基准的单骨骼对比学习方法。多骨骼集成进一步提升性能,在两个数据集上均达到新最佳水平。
原文摘要 · Abstract (English)
Contrastive learning has gained significant attention in skeleton-based action recognition for its ability to learn robust representations from unlabeled data. However, existing methods rely on a single skeleton convention, which limits their ability to generalize across datasets with diverse joint structures and anatomical coverage. We propose Multi-Skeleton Contrastive Learning (MS-CLR), a general self-supervised framework that aligns pose representations across multiple skeleton conventions extracted from the same sequence. This encourages the model to learn structural invariances and capture diverse anatomical cues, resulting in more expressive and generalizable features. To support this, we adapt the ST-GCN architecture to handle skeletons with varying joint layouts and scales through a unified representation scheme. Experiments on the NTU RGB+D 60 and 120 datasets demonstrate that MS-CLR consistently improves performance over strong single-skeleton contrastive learning baselines. A multi-skeleton ensemble further boosts performance, setting new state-of-the-art results on both datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。