融合生理信号与深度模型,提升情绪识别准确率。
Deep Temporal Modeling and Ensemble Fusion for Multimodal Emotion Recognition from Physiological Signals

- 用LSTM、TCN、Transformer处理腕部和胸部信号
- 集成模型达98.91%准确率,优于单一模型
- 适合健康监测与情感计算研究者参考
生理压力与情绪识别对健康监控和情感计算至关重要。本文在WESAD数据集上评估了LSTM、TCN和Transformer等深度学习模型在多模态情感识别中的表现,使用腕部和胸部传感器信号。通过消融实验分析各模态单独贡献,分别训练仅腕部和仅胸部输入的模型。同时采用早期融合(信号级拼接)和晚期集成融合策略,将三类模型在多模态输入下的预测结果合并。结果表明,Transformer在多模态设置中持续表现最佳,而TCN在仅腕部输入时最优。集成方法取得最高整体准确率(98.91 ± 0.13%)和宏F1分数(98.56 ± 0.17%)。这些发现证明了传感器融合与集成融合在构建鲁棒生理情绪识别系统中的有效性。
原文摘要 · Abstract (English)
Physiological stress and emotion recognition are important for health monitoring and affective computing. In this work, we present a comprehensive evaluation of deep learning models such as Long Short-Term Memory (LSTM), Temporal Convolutional Networks (TCN), and Transformer on the WESAD dataset for multimodal affect recognition using wrist and chest sensor signals. We perform ablation studies to assess the individual contributions of each modality by training models on wrist-only and chest-only inputs. In addition, we implement a late-fusion ensemble strategy that combines predictions from all three architectures trained on multimodal input. We also employ early fusion at the sensor level by concatenating wrist and chest signals before feeding them into each model. Our results show that Transformer models consistently achieve the highest accuracy in multimodal settings, while TCN models perform best in the wrist-only configuration. The ensemble method yields the highest overall accuracy (98.91 +/- 0.13%) and macro-F1 score (98.56 +/- 0.17%). These findings demonstrate the effectiveness of sensor fusion and ensemble-based fusion in developing robust systems for physiological emotion recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。