arXiv:2511.14179cs.CV2025-11

用博弈论优化骨骼动作识别的自监督对比学习,提升关键运动区域建模与负样本质量。

DoGCLR: Dominance-Game Contrastive Learning Network for Skeleton-Based Action Recognition

  • 基于博弈论动态构建正负样本,实现语义保留与区分力平衡。
  • 在NTU RGB+D上达到81.1%/89.4%准确率,优于当前最优方法。
  • 适合需要高鲁棒性骨骼动作识别的场景,如复杂动作或数据受限环境。

现有基于骨骼的动作识别自监督对比学习方法通常对所有骨骼区域一视同仁,并采用先进先出队列存储负样本,导致运动信息丢失且负样本选择不佳。为此,本文提出基于博弈论的骨骼动作识别自监督框架DoGCLR,将正负样本构建建模为动态主导博弈,使两类样本交互以达成语义保持与判别强度的平衡。具体而言,时空双重加权定位机制识别关键运动区域,指导区域级增强以提升运动多样性并保持语义;同时,基于熵的主导策略管理内存池,保留高熵(难)负样本并替换低熵(弱)样本,确保持续接收有效对比信号。在NTU RGB+D和PKU-MMD数据集上进行大量实验:在NTU RGB+D 60 X-Sub/X-View上分别取得81.1%/89.4%准确率,在120 X-Sub/X-Set上达到71.2%/75.5%,超越现有方法0.1%、2.7%、1.1%和2.3%。在PKU-MMD Part I/Part II上表现接近最先进水平,且在Part II上高出1.9%,凸显其在更具挑战性场景下的强鲁棒性。

原文摘要 · Abstract (English)

Existing self-supervised contrastive learning methods for skeleton-based action recognition often process all skeleton regions uniformly, and adopt a first-in-first-out (FIFO) queue to store negative samples, which leads to motion information loss and non-optimal negative sample selection. To address these challenges, this paper proposes Dominance-Game Contrastive Learning network for skeleton-based action Recognition (DoGCLR), a self-supervised framework based on game theory. DoGCLR models the construction of positive and negative samples as a dynamic Dominance Game, where both sample types interact to reach an equilibrium that balances semantic preservation and discriminative strength. Specifically, a spatio-temporal dual weight localization mechanism identifies key motion regions and guides region-wise augmentations to enhance motion diversity while maintaining semantics. In parallel, an entropy-driven dominance strategy manages the memory bank by retaining high entropy (hard) negatives and replacing low-entropy (weak) ones, ensuring consistent exposure to informative contrastive signals. Extensive experiments are conducted on NTU RGB+D and PKU-MMD datasets. On NTU RGB+D 60 X-Sub/X-View, DoGCLR achieves 81.1%/89.4% accuracy, and on NTU RGB+D 120 X-Sub/X-Set, DoGCLR achieves 71.2%/75.5% accuracy, surpassing state-of-the-art methods by 0.1%, 2.7%, 1.1%, and 2.3%, respectively. On PKU-MMD Part I/Part II, DoGCLR performs comparably to the state-of-the-art methods and achieves a 1.9% higher accuracy on Part II, highlighting its strong robustness on more challenging scenarios.

骨骼动作识别自监督学习对比学习博弈论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。