用跨模态对比蒸馏提升单模态情绪识别准确率
CMCRD: Cross-Modal Contrastive Representation Distillation for Emotion Recognition
- 用预训练眼动模型辅助脑电模型训练,反向提升特征提取能力
- 仅用脑电或眼动数据测试,准确率比纯脑电模型高6.2%平均
- 适合资源受限但需高精度情绪识别的场景
情绪识别是情感计算和人机交互的重要组成部分。单模态情绪识别虽便捷,但准确率可能不足;多模态识别更准确,却增加数据采集的复杂性和成本。本文提出跨模态对比表示蒸馏(CMCRD),在训练中同时使用脑电(EEG)和眼动数据,但测试时仅需其中一种信号。该方法利用预训练的眼动分类模型辅助脑电模型训练,或反之,增强特征提取能力。实验在三个多模态情绪识别数据集上,采用三种神经网络架构验证,结果表明相比纯脑电模型,平均分类准确率提升约6.2%,既提高精度又简化系统部署。
原文摘要 · Abstract (English)
Emotion recognition is an important component of affective computing, and also human-machine interaction. Unimodal emotion recognition is convenient, but the accuracy may not be high enough; on the contrary, multi-modal emotion recognition may be more accurate, but it also increases the complexity and cost of the data collection system. This paper considers cross-modal emotion recognition, i.e., using both electroencephalography (EEG) and eye movement in training, but only EEG or eye movement in test. We propose cross-modal contrastive representation distillation (CMCRD), which uses a pre-trained eye movement classification model to assist the training of an EEG classification model, improving feature extraction from EEG signals, or vice versa. During test, only EEG signals (or eye movement signals) are acquired, eliminating the need for multi-modal data. CMCRD not only improves the emotion recognition accuracy, but also makes the system more simplified and practical. Experiments using three different neural network architectures on three multi-modal emotion recognition datasets demonstrated the effectiveness of CMCRD. Compared with the EEG-only model, it improved the average classification accuracy by about 6.2%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。