用脑电与语音联合训练,让系统在缺脑电时仍能准确识情绪。
Unifying EEG and Speech for Emotion Recognition: A Two-Step Joint Learning Framework for Handling Missing EEG Data During Inference
- 分两步学习:先单独学语音和脑电特征,再融合成共同表示
- 在无脑电数据时,识别准确率仍超基线模型15%以上
- 适合需要高可靠情绪识别的智能交互场景
计算机界面正向多模态发展以提升人机交互体验。自动情绪识别(AER)可使交互更自然,但语音易被人为伪造,而脑电虽可靠却需特殊设备难以实用。本文提出一种两阶段联合多模态学习框架(JMML),利用训练时的双模态数据,在推理时即使缺少脑电也能保持高精度情绪识别。首先通过JEC-SSL对各模态独立进行内部学习;随后采用改进的深度典型相关交叉自编码器(E-DCC-CAE)实现跨模态对齐,将语音与脑电映射至共享表示空间以最大化相关性。由此生成的情绪嵌入融合了双模态特性,显著提升分类器性能。实验表明该方法有效,是首个将语音与脑电结合用于可靠AER的联合学习尝试。
原文摘要 · Abstract (English)
Computer interfaces are advancing towards using multi-modalities to enable better human-computer interactions. The use of automatic emotion recognition (AER) can make the interactions natural and meaningful thereby enhancing the user experience. Though speech is the most direct and intuitive modality for AER, it is not reliable because it can be intentionally faked by humans. On the other hand, physiological modalities like EEG, are more reliable and impossible to fake. However, use of EEG is infeasible for realistic scenarios usage because of the need for specialized recording setup. In this paper, one of our primary aims is to ride on the reliability of the EEG modality to facilitate robust AER on the speech modality. Our approach uses both the modalities during training to reliably identify emotion at the time of inference, even in the absence of the more reliable EEG modality. We propose, a two-step joint multi-modal learning approach (JMML) that exploits both the intra- and inter- modal characteristics to construct emotion embeddings that enrich the performance of AER. In the first step, using JEC-SSL, intra-modal learning is done independently on the individual modalities. This is followed by an inter-modal learning using the proposed extended variant of deep canonically correlated cross-modal autoencoder (E-DCC-CAE). The approach learns the joint properties of both the modalities by mapping them into a common representation space, such that the modalities are maximally correlated. These emotion embeddings, hold properties of both the modalities there by enhancing the performance of ML classifier used for AER. Experimental results show the efficacy of the proposed approach. To best of our knowledge, this is the first attempt to combine speech and EEG with joint multi-modal learning approach for reliable AER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。