用混合模型提升阿拉伯语情感识别准确率
Hybrid CNN-Transformer Architecture for Arabic Speech Emotion Recognition

- 结合卷积层与注意力机制,提取语音频谱特征和时序依赖
- 在埃及阿拉伯语数据集上达97.8%准确率,宏F1为0.98
- 为低资源语言的情感识别提供有效技术路径
基于机器学习的情感语音识别在构建以人为本的应用中日益重要。尽管英语、德语等欧洲及亚洲语言已有大量研究,但阿拉伯语相关研究仍较少,主要受限于标注数据集稀缺。本文提出一种基于混合卷积-注意力架构的阿拉伯语语音情感识别(SER)系统。模型利用卷积层从梅尔频谱图输入中提取判别性频谱特征,同时通过Transformer编码器捕捉语音中的长程时序依赖。实验在EYASE(埃及阿拉伯语语音情感)语料库上进行,所提模型达到97.8%的准确率和0.98的宏平均F1分数。结果表明,结合卷积特征提取与注意力建模在阿拉伯语情感识别中有效,凸显了基于Transformer的方法在低资源语言中的潜力。
原文摘要 · Abstract (English)
Recognizing emotions from speech using machine learning has become an active research area due to its importance in building human-centered applications. However, while many studies have been conducted in English, German, and other European and Asian languages, research in Arabic remains scarce because of the limited availability of annotated datasets. In this paper, we present an Arabic Speech Emotion Recognition (SER) system based on a hybrid CNN-Transformer architecture. The model leverages convolutional layers to extract discriminative spectral features from Mel-spectrogram inputs and Transformer encoders to capture long-range temporal dependencies in speech. Experiments were conducted on the EYASE (Egyptian Arabic speech emotion) corpus, and the proposed model achieved 97.8% accuracy and a macro F1-score of 0.98. These results demonstrate the effectiveness of combining convolutional feature extraction with attention-based modeling for Arabic SER and highlight the potential of Transformer-based approaches in low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。