arXiv:2606.13030cs.CV2026-06被引 1

通过伪标签与语义对齐提升微手势识别跨人泛化能力

A Multi-Modal Framework with Cross-Subject Pseudo-Labeling and Semantic Alignment for Micro-Gesture Recognition

论文配图:A Multi-Modal Framework with Cross-Subject Pseudo-Labeling and Semantic Alignment for Micro-Gesture Recognition
图 1 · 摘自论文原文
  • 融合骨骼、热力图与视觉特征,提取细粒度动作表征
  • 在未修剪视频中实现68.13%的F1分数,显著提升尾部类识别
  • 适合跨被试微手势识别研究者参考

微手势(MGs)是频繁表达隐性情绪的自发细微动作。在未修剪视频中识别微手势仍极具挑战,因其信号噪声比极低、类别分布严重长尾,且跨被试评估存在固有领域偏移。本文提出针对第四届MiGA-IJCAI挑战赛第1赛道的多模态综合框架。为捕捉细粒度表征,设计了基于显著性引导的多模态提取流程,融合68关键点骨骼坐标、3D热力图体积与高分辨率RGB视觉特征。引入温和的平方根平滑加权机制,结合正交语义嵌入损失,保护尾部类别而不影响整体性能。更重要的是,提出跨模态伪标签(CMPL)策略用于无监督域适应,显著增强单模态鲁棒性。最后采用温度缩放软投票机制缓解融合阶段过自信问题。大量实验表明,本框架取得68.13%的竞争力F1分数,位列第4。

原文摘要 · Abstract (English)

Micro-gestures (MGs) are spontaneous and subtle body movements that frequently convey hidden human emotions. Recognizing MGs in untrimmed videos remains highly challenging due to their extremely low signal-to-noise ratio, severe long-tailed class distribution, and the inherent domain shift encountered in cross-subject evaluation scenarios. In this paper, we propose a comprehensive multi-modal framework for Track 1 of the 4th MiGA-IJCAI Challenge. To capture fine-grained representations, we design a saliency-guided multi-modal extraction pipeline integrating 68-keypoint skeleton joint coordinates, 3D heatmap volumes, and high-resolution RGB visual features. We introduce a gentle square-root smoothed weighting mechanism paired with an Orthogonal Semantic Embedding Loss to protect tail classes without compromising overall recognition capabilities. More importantly, to bridge the cross-subject generalization gap, we propose a Cross-Modal Pseudo-Labeling (CMPL) strategy for unsupervised domain adaptation, which significantly boosts single-modal robustness. A temperature-scaled soft-voting mechanism is finally utilized to alleviate overconfidence during late fusion. Extensive experiments demonstrate that our framework achieves a competitive F1-score of 68.13\%, securing the 4th place.

微手势识别多模态学习伪标签跨被试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。