用对比学习提升少样本下多模态动作识别效果
Contrastive Learning for Multimodal Human Activity Recognition with Limited Labeled Data

- 两阶段训练:先共享特征,再融合特定特征
- 在三个数据集上准确率超越现有方法,收敛更快
- 适合标签稀缺的智能穿戴设备动作识别场景
人类活动识别是众多新兴应用的基础。近年来,研究者利用多源传感器协同感知复杂动态的人类行为。然而,多模态人体活动感知通常面临模态间数据高度异构及标签稀缺的问题,导致现有方案与真实需求之间存在应用鸿沟。本文提出一种通用的对比学习框架CLMM,可在有限标注数据下实现有效的多模态活动识别。CLMM采用新颖的两阶段训练策略:第一阶段使用CNN-DiffTransformer编码器提取局部与全局特征,通过难正样本加权算法增强梯度传播,强化共享学习;第二阶段采用双分支结构,结合质量引导注意力和双向门控单元捕捉模态特异性信息,并通过主-辅协同训练策略融合共享与特异性信息。在三个公开数据集上的实验结果表明,CLMM显著提升了当前最优基线在识别准确率和收敛性能上的表现。
原文摘要 · Abstract (English)
Human activity recognition serves as the foundation for various emerging applications. In recent years, researchers have used collaborative sensing of multi-source sensors to capture complex and dynamic human activities. However, multimodal human activity sensing typically encounters highly heterogeneous data across modalities and label scarcity, resulting in an application gap between existing solutions and real-world needs. In this paper, we propose CLMM, a general contrastive learning framework for human activity recognition that achieves effective multimodal recognition with limited labeled data. CLMM employs a novel two-stage training strategy. In the first stage, CLMM employs a CNN-DiffTransformer encoder to capture cross-modal shared information by extracting local and global features. Meanwhile, a hard-positive samples weighting algorithm enhances gradient propagation to reinforce shared learning. In the second stage, a dual-branch architecture combining quality-guided attention and bidirectional gated units captures modality-specific information, while a primary-auxiliary collaborative training strategy fuses both shared and modality-specific information. Experimental results on three public datasets demonstrate that CLMM significantly improves state-of-the-art baselines in both recognition accuracy and convergence performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。