arXiv:2412.09044cs.CVcs.AI2024-12中稿 · AAAI被引 2

通过结构与步态特征提升骨骼表征,显著增强人体重识别效果。

Motif Guided Graph Transformer with Combinatorial Skeleton Prototype Learning for Skeleton-Based Person Re-Identification

  • 引入层级结构与步态协同的图注意力机制,捕捉关节局部关联与关键肢体协作。
  • 设计组合原型学习,通过随机时空组合生成多样子骨架表示,提升判别力。
  • 适用于多种骨骼输入与无监督场景,通用性强,性能超越现有方法。

基于3D骨骼的人体重识别在诸多场景中具有重要价值,但挑战巨大。现有方法通常假设所有关节间存在虚拟运动关系,并采用平均关节或序列表示进行学习,却忽视了关键身体结构与步态特征,难以充分挖掘骨骼数据中的时空子模式。本文提出通用的动机引导图变压器与组合骨架原型学习框架(MoCos),利用结构特异性和步态相关性建模关节关系,并通过骨架图的随机时空组合生成多样化子骨架与子轨迹表示,与每类身份的代表性特征(原型)进行对比,学习类别相关语义与判别性表征。实验表明,MoCos在多个基准上优于现有最优模型,且在RGB估计骨骼、不同图建模及无监督场景下均表现良好,具备强泛化能力。

原文摘要 · Abstract (English)

Person re-identification (re-ID) via 3D skeleton data is a challenging task with significant value in many scenarios. Existing skeleton-based methods typically assume virtual motion relations between all joints, and adopt average joint or sequence representations for learning. However, they rarely explore key body structure and motion such as gait to focus on more important body joints or limbs, while lacking the ability to fully mine valuable spatial-temporal sub-patterns of skeletons to enhance model learning. This paper presents a generic Motif guided graph transformer with Combinatorial skeleton prototype learning (MoCos) that exploits structure-specific and gait-related body relations as well as combinatorial features of skeleton graphs to learn effective skeleton representations for person re-ID. In particular, motivated by the locality within joints' structure and the body-component collaboration in gait, we first propose the motif guided graph transformer (MGT) that incorporates hierarchical structural motifs and gait collaborative motifs, which simultaneously focuses on multi-order local joint correlations and key cooperative body parts to enhance skeleton relation learning. Then, we devise the combinatorial skeleton prototype learning (CSP) that leverages random spatial-temporal combinations of joint nodes and skeleton graphs to generate diverse sub-skeleton and sub-tracklet representations, which are contrasted with the most representative features (prototypes) of each identity to learn class-related semantics and discriminative skeleton representations. Extensive experiments validate the superior performance of MoCos over existing state-of-the-art models. We further show its generality under RGB-estimated skeletons, different graph modeling, and unsupervised scenarios.

骨骼重识别图神经网络原型学习动作识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。