轻量模型 ArabEmoNet 提升阿拉伯语语音情感识别准确率
ArabEmoNet: A Lightweight Hybrid 2D CNN-BiLSTM Model with Attention for Robust Arabic Speech Emotion Recognition
- 用2D卷积处理梅尔频谱图,保留更多情绪细节
- 仅100万参数,比HuBERT小90倍仍性能领先
- 适合低资源环境,推动真实场景应用
语音情感识别对人机交互至关重要,尤其对于数据稀缺的阿拉伯语等低资源语言。本文提出 ArabEmoNet,一种轻量级混合2D CNN-BiLSTM模型,通过2D卷积处理梅尔频谱图,捕捉传统方法易丢失的时频动态特征。相比依赖离散MFCC与1D卷积的旧系统,该模型在保持高性能的同时大幅降低复杂度:仅100万参数,比HuBERT base小90倍,比Whisper小74倍,实现更优的准确率与推理效率,显著提升阿拉伯语语音情感识别的实用性与可部署性。
原文摘要 · Abstract (English)
Speech emotion recognition is vital for human-computer interaction, particularly for low-resource languages like Arabic, which face challenges due to limited data and research. We introduce ArabEmoNet, a lightweight architecture designed to overcome these limitations and deliver state-of-the-art performance. Unlike previous systems relying on discrete MFCC features and 1D convolutions, which miss nuanced spectro-temporal patterns, ArabEmoNet uses Mel spectrograms processed through 2D convolutions, preserving critical emotional cues often lost in traditional methods. While recent models favor large-scale architectures with millions of parameters, ArabEmoNet achieves superior results with just 1 million parameters, 90 times smaller than HuBERT base and 74 times smaller than Whisper. This efficiency makes it ideal for resource-constrained environments. ArabEmoNet advances Arabic speech emotion recognition, offering exceptional performance and accessibility for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。