轻量级多尺度模型提升语音情感识别准确率与泛化能力
SETEAB: Multiscale approach with Squeeze-and-Excitation Temporal Enhanced Aware Block for Speech Emotion Recognition

- 采用深度卷积下采样减少计算量,保留关键情感特征
- 引入通道注意力机制增强特征鲁棒性,提升识别精度
- 设计时序增强模块捕捉长程依赖,适合跨数据集应用
本文提出一种新型轻量级多尺度架构用于语音情感识别(SER),包含三项创新:首先,引入基于深度卷积的下采样模块,在降低模型规模和计算复杂度的同时保留显著情感线索;其次,集成挤压-激励模块,增强通道级重校准,提升表示鲁棒性;第三,设计新型时序增强感知模块,强化时序依赖建模,生成更具区分性的情感感知特征。该模型旨在协同提升紧凑性、识别性能与泛化能力。在基准SER数据集上的实验表明,该方法在降低计算复杂度的同时实现更高准确率,并且相比多数近期先进网络在跨语料库任务上表现更优。
原文摘要 · Abstract (English)
This paper proposes a novel lightweight multiscale architecture for speech emotion recognition (SER) with three key innovations. First, a depthwise convolution-based subsampling module is introduced to reduce model size and computation while preserving salient emotional cues. Second, a Squeeze-and-Excitation block is integrated to enhance channel-wise recalibration and improve representation robustness. Third, a new Temporal Enhanced Aware Block is designed to strengthen temporal dependency modeling and produce more discriminative emotion-aware features. The proposed model is explicitly designed to jointly improve compactness, recognition performance, and generalizability. Experiments on benchmark SER datasets show that our method achieves higher accuracy with reduced computational complexity, while also delivering stronger cross-corpus performance than most recent advanced networks for SER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。