用专家路由联合提升语音增强与情绪识别的抗噪能力
Joint Learning using Mixture-of-Expert-Based Representation for Speech Enhancement and Robust Emotion Recognition
- 帧级专家路由选择自监督表示,动态适配任务需求
- 在-5dB信噪比下情绪识别F1提升12.0%,语音增强SSNR提升28.2%
- 适合需要低延迟、高鲁棒性的实时情感交互系统
语音情绪识别(SER)在构建情感感知语音系统中至关重要,但在噪声环境下性能显著下降。尽管语音增强(SE)可提升鲁棒性,但常引入失真并遮蔽情绪线索,且增加计算开销。多任务学习(MTL)提供了联合优化的替代方案,但传统共享主干模型常因梯度干扰和表征冲突导致性能受限。为此,本文提出稀疏专家混合表示集成技术(Sparse MERIT),一种基于自监督语音表示的灵活MTL框架,通过帧级专家路由实现动态任务适配。该方法引入任务特异性门控网络,从共享专家池中为每帧选择最优表示,实现参数高效且任务自适应的学习。在MSP-Podcast数据集上的实验表明,Sparse MERIT在两种任务上均优于基线模型。在最严苛的-5 dB信噪比条件下,其情绪识别F1-macro平均比依赖预处理增强的基线提升12.0%,比朴素MTL基线提升3.4%,且在未见噪声条件下具有统计显著性;语音增强方面,相比预处理基线,段落信噪比(SSNR)提升28.2%,比朴素MTL基线提升20.0%。结果证明Sparse MERIT在噪声环境中对情绪识别与增强任务均具备强鲁棒性和泛化能力。
原文摘要 · Abstract (English)
Speech emotion recognition (SER) plays a critical role in building emotion-aware speech systems, but its performance degrades significantly under noisy conditions. Although speech enhancement (SE) can improve robustness, it often introduces artifacts that obscure emotional cues and adds computational overhead to the pipeline. Multi-task learning (MTL) offers an alternative by jointly optimizing SE and SER tasks. However, conventional shared-backbone models frequently suffer from gradient interference and representational conflicts between tasks. To address these challenges, we propose the Sparse Mixture-of-Experts Representation Integration Technique (Sparse MERIT), a flexible MTL framework that applies frame-wise expert routing over self-supervised speech representations. Sparse MERIT incorporates task-specific gating networks that dynamically select from a shared pool of experts for each frame, enabling parameter-efficient and task-adaptive representation learning. Experiments on the MSP-Podcast corpus show that Sparse MERIT consistently outperforms baseline models on both SER and SE tasks. Under the most challenging condition of -5 dB signal-to-noise ratio (SNR), Sparse MERIT improves SER F1-macro by an average of 12.0% over a baseline relying on a SE pre-processing strategy, and by 3.4% over a naive MTL baseline, with statistical significance on unseen noise conditions. For SE, Sparse MERIT improves segmental SNR (SSNR) by 28.2% over the SE pre-processing baseline and by 20.0% over the naive MTL baseline. These results demonstrate that Sparse MERIT provides robust and generalizable performance for both emotion recognition and enhancement tasks in noisy environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。