用骨骼姿态引导视觉语言模型,提升微手势识别准确率。
CLIP-MG: Guiding Semantic Attention with Skeletal Pose Features and RGB Data for Micro-Gesture Recognition on the iMiGUE Dataset
- 结合骨骼姿态生成语义查询,实现动作特征精准定位
- 在iMiGUE数据集上达到61.82%的分类准确率
- 适合关注微手势识别与多模态融合的研究者
微手势识别在情感计算中极具挑战性,因其动作细微、非自主且运动幅度极小。本文提出一种基于CLIP的微手势识别架构——CLIP-MG,针对iMiGUE数据集进行定制优化。该模型通过骨骼姿态引导的语义查询生成和门控多模态融合机制,将人体姿态信息融入视觉-语言识别流程。实验结果表明,该模型在测试集上取得61.82%的Top-1准确率,验证了该方法的有效性,同时也揭示了当前视觉-语言模型在微手势识别任务中仍面临显著困难。
原文摘要 · Abstract (English)
Micro-gesture recognition is a challenging task in affective computing due to the subtle, involuntary nature of the gestures and their low movement amplitude. In this paper, we introduce a Pose-Guided Semantics-Aware CLIP-based architecture, or CLIP for Micro-Gesture recognition (CLIP-MG), a modified CLIP model tailored for micro-gesture classification on the iMiGUE dataset. CLIP-MG integrates human pose (skeleton) information into the CLIP-based recognition pipeline through pose-guided semantic query generation and a gated multi-modal fusion mechanism. The proposed model achieves a Top-1 accuracy of 61.82%. These results demonstrate both the potential of our approach and the remaining difficulty in fully adapting vision-language models like CLIP for micro-gesture recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。