arXiv:2409.10687eess.AScs.HC2024-09中稿 · ICRA被引 5

用视觉变压器提升人机交互中的个性化语音情绪识别

Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers

  • 用ViT和BEiT模型微调捕捉个体语音特征
  • 集成模型在四类情绪识别上准确率最高
  • 适合需要个性化情感理解的机器人场景

情绪是言语交流的关键要素,因此在人机交互(HRI)中理解个体情感至关重要。本文研究了视觉变压器模型(ViT 和 BEiT)在人机交互语音情绪识别(SER)中的应用。通过在基准数据集上微调这些模型,并采用集成方法,使模型能更好地适应个体语音特征。为此,我们收集了不同参与者与NAO机器人进行伪自然对话时的音频数据。随后对基于ViT和BEiT的模型进行微调,并在未见语音样本上测试其表现。结果表明,在识别四种基本情绪(中性、快乐、悲伤、愤怒)时,先在基准数据集上微调再使用或集成ViT/BEiT模型,比直接微调原始模型获得更高的个体分类准确率。

原文摘要 · Abstract (English)

Emotions are an essential element in verbal communication, so understanding individuals' affect during a human-robot interaction (HRI) becomes imperative. This paper investigates the application of vision transformer models, namely ViT (Vision Transformers) and BEiT (BERT Pre-Training of Image Transformers) pipelines, for Speech Emotion Recognition (SER) in HRI. The focus is to generalize the SER models for individual speech characteristics by fine-tuning these models on benchmark datasets and exploiting ensemble methods. For this purpose, we collected audio data from different human subjects having pseudo-naturalistic conversations with the NAO robot. We then fine-tuned our ViT and BEiT-based models and tested these models on unseen speech samples from the participants. In the results, we show that fine-tuning vision transformers on benchmark datasets and and then using either these already fine-tuned models or ensembling ViT/BEiT models gets us the highest classification accuracies per individual when it comes to identifying four primary emotions from their speech: neutral, happy, sad, and angry, as compared to fine-tuning vanilla-ViTs or BEiTs.

语音识别视觉变压器人机交互情绪识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。