arXiv:2411.00822cs.CVcs.AI2024-11被引 16

融合脑电与多模态数据,提升情绪识别准确率。

EEG-based Multimodal Representation Learning for Emotion Recognition

  • 设计可动态调整注意力的多模态框架,适配不同输入大小。
  • 在三模态情感数据集上实现优于基线的方法性能。
  • 适合研究脑电融合与跨模态学习的科研人员参考。

多模态学习是当前研究热点,但将脑电图(EEG)数据融入其中面临固有变异性与数据稀缺的挑战。本文提出一种新型多模态框架,不仅支持视频、图像、音频等传统模态,还整合了EEG数据。该框架能灵活处理不同输入尺寸,并动态调节注意力以反映各模态特征的重要性。我们在一个新近发布的三模态情感识别数据集上评估方法,该数据集包含视频、音频和EEG数据,为多模态学习提供了理想的测试平台。实验结果为该数据集建立了基准性能,并验证了所提框架的有效性。本工作展示了将EEG融入多模态系统的优势,为情绪识别及其他应用提供了更鲁棒、更全面的技术路径。

原文摘要 · Abstract (English)

Multimodal learning has been a popular area of research, yet integrating electroencephalogram (EEG) data poses unique challenges due to its inherent variability and limited availability. In this paper, we introduce a novel multimodal framework that accommodates not only conventional modalities such as video, images, and audio, but also incorporates EEG data. Our framework is designed to flexibly handle varying input sizes, while dynamically adjusting attention to account for feature importance across modalities. We evaluate our approach on a recently introduced emotion recognition dataset that combines data from three modalities, making it an ideal testbed for multimodal learning. The experimental results provide a benchmark for the dataset and demonstrate the effectiveness of the proposed framework. This work highlights the potential of integrating EEG into multimodal systems, paving the way for more robust and comprehensive applications in emotion recognition and beyond.

情绪识别脑电融合多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。