arXiv:2409.14221eess.AScs.SD2024-09被引 3

多模态模型结合最优传输,提升非语言情绪识别准确率

Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition

  • 用最优传输注意力对齐多模态表示,融合语言与图像绑定模型
  • 在三个基准数据集上达到76.47%准确率,刷新非语言情绪识别纪录
  • 适合做多模态情感分析、语音情绪识别的研究者参考

本研究探讨了多模态基础模型(MFMs)在非语言声音情绪识别(NVER)中的应用。我们假设,由于跨模态联合预训练,MFMs能更有效地解析和区分音频中细微的情绪线索,从而优于仅基于音频的基础模型(AFMs)。为此,我们提取了当前最先进的MFMs与AFMs的表征,并在标准的NVER数据集上进行评估。同时,受语音识别和音频深度伪造检测研究启发,我们探索了融合多个基础模型表征以进一步提升性能的可能性。为此,提出一种名为MATA(模内对齐通过传输注意力)的框架。结合语言绑定(LanguageBind)与图像绑定(ImageBind)模型,利用MATA实现了最高性能:在ASVP-ESD、JNV、VIVAE数据集上的准确率分别为76.47%、77.40%、75.12%,F1分数分别为70.35%、76.19%、74.63%,显著优于单一模型及基线融合方法,在基准数据集上达到最先进水平。

原文摘要 · Abstract (English)

In this study, we investigate multimodal foundation models (MFMs) for emotion recognition from non-verbal sounds. We hypothesize that MFMs, with their joint pre-training across multiple modalities, will be more effective in non-verbal sounds emotion recognition (NVER) by better interpreting and differentiating subtle emotional cues that may be ambiguous in audio-only foundation models (AFMs). To validate our hypothesis, we extract representations from state-of-the-art (SOTA) MFMs and AFMs and evaluated them on benchmark NVER datasets. We also investigate the potential of combining selected foundation model representations to enhance NVER further inspired by research in speech recognition and audio deepfake detection. To achieve this, we propose a framework called MATA (Intra-Modality Alignment through Transport Attention). Through MATA coupled with the combination of MFMs: LanguageBind and ImageBind, we report the topmost performance with accuracies of 76.47%, 77.40%, 75.12% and F1-scores of 70.35%, 76.19%, 74.63% for ASVP-ESD, JNV, and VIVAE datasets against individual FMs and baseline fusion techniques and report SOTA on the benchmark datasets.

情绪识别多模态最优传输基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。