arXiv:2509.00077eess.AScs.AI2025-09被引 3

用数据增强和迁移学习,小样本下提升语音情感识别准确率

Amplifying Emotional Signals: Data-Efficient Deep Learning for Robust Speech Emotion Recognition

  • 结合预训练模型与创新数据增强策略
  • 在小数据集上达到66.7%准确率和0.631 F1分数
  • 适合资源有限但需高鲁棒性的语音交互场景

语音情感识别(SER)在人机交互中面临持续挑战。尽管深度学习推动了语音处理发展,但在小规模数据集上实现高性能仍是关键难题。本文评估了支持向量机(SVM)、长短期记忆网络(LSTM)和卷积神经网络(CNN)等模型在语音情感分类中的表现。通过战略性运用迁移学习和创新的数据增强技术,模型在受限数据条件下仍取得显著性能提升。最有效模型为ResNet34架构,在合并的RAVDESS与SAVEE数据集上达到66.7%准确率和0.631 F1分数,创下新基准。结果表明,利用预训练模型与数据增强能有效缓解数据稀缺问题,推动更鲁棒、泛化能力强的SER系统发展。

原文摘要 · Abstract (English)

Speech Emotion Recognition (SER) presents a significant yet persistent challenge in human-computer interaction. While deep learning has advanced spoken language processing, achieving high performance on limited datasets remains a critical hurdle. This paper confronts this issue by developing and evaluating a suite of machine learning models, including Support Vector Machines (SVMs), Long Short-Term Memory networks (LSTMs), and Convolutional Neural Networks (CNNs), for automated emotion classification in human speech. We demonstrate that by strategically employing transfer learning and innovative data augmentation techniques, our models can achieve impressive performance despite the constraints of a relatively small dataset. Our most effective model, a ResNet34 architecture, establishes a new performance benchmark on the combined RAVDESS and SAVEE datasets, attaining an accuracy of 66.7% and an F1 score of 0.631. These results underscore the substantial benefits of leveraging pre-trained models and data augmentation to overcome data scarcity, thereby paving the way for more robust and generalizable SER systems.

语音情感识别小样本学习数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。