通过博弈论框架提升骨骼动作识别的自监督学习效果
M3GCLR: Multi-View Mini-Max Infinite Skeleton-Data Game Contrastive Learning For Skeleton-Based Action Recognition
- 构建多视角对抗性游戏模型,实现最小最大优化
- 在NTU RGB+D上达到85.8%准确率,超越当前最优
- 适合研究自监督学习与动作识别的学者参考
近年来,对比学习因其减少对标注数据依赖的能力受到广泛关注。然而,现有的自监督骨骼动作识别方法仍存在三大问题:视角差异建模不足、缺乏有效对抗机制、增强扰动不可控。为此,我们提出多视角极小极大无限骨骼数据博弈对比学习(M3GCLR),一种基于博弈论的对比学习框架。首先,建立无限骨骼数据博弈(ISG)模型及等价定理,并提供严格证明,支持基于多视角互信息的极小极大优化。其次,通过多视角旋转增强生成正常-极端数据对,采用时序平均输入作为中立锚点,显式刻画扰动强度。再次,基于所提等价定理,构建强对抗性的极小极大骨骼数据博弈,促使模型挖掘更丰富的动作判别信息。最后,引入双损失等价优化器以优化博弈平衡,使学习过程最大化动作相关信息同时最小化编码冗余,并证明其与ISG模型等价。大量实验表明,M3GCLR在NTU RGB+D 60(X-Sub, X-View)上分别取得82.1%、85.8%准确率,在NTU RGB+D 120(X-Sub, X-Set)上达72.3%、75.0%,在PKU-MMD Part I、II三流设置下分别获得89.1%、45.2%,均达到或超过现有最优性能。消融实验验证了各组件的有效性。
原文摘要 · Abstract (English)
In recent years, contrastive learning has drawn significant attention as an effective approach to reducing reliance on labeled data. However, existing methods for self-supervised skeleton-based action recognition still face three major limitations: insufficient modeling of view discrepancies, lack of effective adversarial mechanisms, and uncontrollable augmentation perturbations. To tackle these issues, we propose the Multi-view Mini-Max infinite skeleton-data Game Contrastive Learning for skeleton-based action Recognition (M3GCLR), a game-theoretic contrastive framework. First, we establish the Infinite Skeleton-data Game (ISG) model and the ISG equilibrium theorem, and further provide a rigorous proof, enabling mini-max optimization based on multi-view mutual information. Then, we generate normal-extreme data pairs through multi-view rotation augmentation and adopt temporally averaged input as a neutral anchor to achieve structural alignment, thereby explicitly characterizing perturbation strength. Next, leveraging the proposed equilibrium theorem, we construct a strongly adversarial mini-max skeleton-data game to encourage the model to mine richer action-discriminative information. Finally, we introduce the dual-loss equilibrium optimizer to optimize the game equilibrium, allowing the learning process to maximize action-relevant information while minimizing encoding redundancy, and we prove the equivalence between the proposed optimizer and the ISG model. Extensive Experiments show that M3GCLR achieves three-stream 82.1%, 85.8% accuracy on NTU RGB+D 60 (X-Sub, X-View) and 72.3%, 75.0% accuracy on NTU RGB+D 120 (X-Sub, X-Set). On PKU-MMD Part I and II, it attains 89.1%, 45.2% in three-stream respectively, all results matching or outperforming state-of-the-art performance. Ablation studies confirm the effectiveness of each component.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。