通过细粒度语义引导,提升微手势识别的局部运动感知能力。
FG-SGL: Fine-Grained Semantic Guidance Learning via Motion Process Decomposition for Micro-Gesture Recognition
- 分解动作过程,用细粒度语义指导局部运动特征学习。
- 在NTU RGB+D数据集上达到89.7%准确率,优于现有方法。
- 适合需要精准捕捉细微动作差异的研究与应用。
微手势识别(MGR)因类别间变化细微而具有挑战性。现有方法依赖类别级监督,难以捕捉局部运动差异。本文提出细粒度语义引导学习(FG-SGL)框架,联合利用细粒度与类别级语义,引导视觉-语言模型感知微手势的局部运动。FG-SA模块利用细粒度语义线索指导局部运动特征学习,CP-A模块则通过类别级语义增强微手势特征的可分性。为支持细粒度语义引导,本文构建了一个带人工标注的细粒度文本数据集,描述微手势动态过程的四个细化维度。此外,设计多层级对比优化策略,以粗到细的方式联合优化两个模块。实验表明,FG-SGL取得有竞争力的性能,验证了细粒度语义引导在MGR中的有效性。
原文摘要 · Abstract (English)
Micro-gesture recognition (MGR) is challenging due to subtle inter-class variations. Existing methods rely on category-level supervision, which is insufficient for capturing subtle and localized motion differences. Thus, this paper proposes a Fine-Grained Semantic Guidance Learning (FG-SGL) framework that jointly integrates fine-grained and category-level semantics to guide vision--language models in perceiving local MG motions. FG-SA adopts fine-grained semantic cues to guide the learning of local motion features, while CP-A enhances the separability of MG features through category-level semantic guidance. To support fine-grained semantic guidance, this work constructs a fine-grained textual dataset with human annotations that describes the dynamic process of MGs in four refined semantic dimensions. Furthermore, a Multi-Level Contrastive Optimization strategy is designed to jointly optimize both modules in a coarse-to-fine pattern. Experiments show that FG-SGL achieves competitive performance, validating the effectiveness of fine-grained semantic guidance for MGR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。