arXiv:2409.05148cs.SDcs.CL2024-09ECCV被引 3

用注意力机制提升西班牙语语音情绪识别准确率

Better Spanish Emotion Recognition In-the-wild: Bringing Attention to Deep Spectrum Voice Analysis

  • 将语音转为图像后,用预训练CNN+注意力机制分类
  • 在两个西班牙语数据集上均超越现有最优模型
  • 验证了模型跨数据集泛化能力,适合真实场景应用

针对社交助手机器人的情感识别需求,本研究聚焦西班牙语语音数据集ELRA-S0329和EmoMatchSpanishDB的副语言特征分析。提出基于DeepSpectrum方法,将音频转换为视觉表示并输入预训练CNN模型,结合支持向量机或全连接神经网络进行分类。本文进一步设计了基于注意力机制的新分类器DS-AM,对比现有最优(SOTA)模型及DeepSpectrum架构,在两个数据集上均取得更好性能。此外,通过跨数据集训练与测试,评估模型对特定数据的依赖性,验证其在真实环境下的泛化能力。

原文摘要 · Abstract (English)

Within the context of creating new Socially Assistive Robots, emotion recognition has become a key development factor, as it allows the robot to adapt to the user's emotional state in the wild. In this work, we focused on the analysis of two voice recording Spanish datasets: ELRA-S0329 and EmoMatchSpanishDB. Specifically, we centered our work in the paralanguage, e.~g. the vocal characteristics that go along with the message and clarifies the meaning. We proposed the use of the DeepSpectrum method, which consists of extracting a visual representation of the audio tracks and feeding them to a pretrained CNN model. For the classification task, DeepSpectrum is often paired with a Support Vector Classifier --DS-SVC--, or a Fully-Connected deep-learning classifier --DS-FC--. We compared the results of the DS-SVC and DS-FC architectures with the state-of-the-art (SOTA) for ELRA-S0329 and EmoMatchSpanishDB. Moreover, we proposed our own classifier based upon Attention Mechanisms, namely DS-AM. We trained all models against both datasets, and we found that our DS-AM model outperforms the SOTA models for the datasets and the SOTA DeepSpectrum architectures. Finally, we trained our DS-AM model in one dataset and tested it in the other, to simulate real-world conditions on how biased is the model to the dataset.

情绪识别语音分析注意力机制西班牙语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。