arXiv:2412.20707eess.AScs.SD2024-12中稿 · ICASSP2025被引 7

利用元数据提升语音情绪识别,通过双阶段微调增强模型表现

Metadata-Enhanced Speech Emotion Recognition: Augmented Residual Integration and Co-Attention in Two-Stage Fine-Tuning

  • 引入增广残差融合模块,保留多层级声学特征
  • 结合协同注意力机制,有效利用元数据上下文关系
  • 在IEMOCAP上超越现有SOTA,适配多种预训练模型

语音情绪识别(SER)旨在分析语音表达以判断说话人的情绪状态,充分且全面地利用音频信息至关重要。为此,我们提出一种基于自监督学习(SSL)模型的新方法,充分利用所有可用的辅助信息——即元数据,以提升性能。通过多任务学习中的双阶段微调策略,我们引入了增广残差融合(ARI)模块,该模块增强了SSL模型编码器中Transformer层的表现,能够高效保留不同层次的声学特征,从而显著提升依赖多层级特征的元数据相关辅助任务性能。此外,由于与ARI具有互补性,协同注意力模块被引入,使模型能有效利用来自元数据辅助任务的多维信息和上下文关系。在预训练基础模型及说话人无关设置下,我们的方法在多个SSL编码器上均一致优于IEMOCAP数据集上的现有最先进模型。

原文摘要 · Abstract (English)

Speech Emotion Recognition (SER) involves analyzing vocal expressions to determine the emotional state of speakers, where the comprehensive and thorough utilization of audio information is paramount. Therefore, we propose a novel approach on self-supervised learning (SSL) models that employs all available auxiliary information -- specifically metadata -- to enhance performance. Through a two-stage fine-tuning method in multi-task learning, we introduce the Augmented Residual Integration (ARI) module, which enhances transformer layers in encoder of SSL models. The module efficiently preserves acoustic features across all different levels, thereby significantly improving the performance of metadata-related auxiliary tasks that require various levels of features. Moreover, the Co-attention module is incorporated due to its complementary nature with ARI, enabling the model to effectively utilize multidimensional information and contextual relationships from metadata-related auxiliary tasks. Under pre-trained base models and speaker-independent setup, our approach consistently surpasses state-of-the-art (SOTA) models on multiple SSL encoders for the IEMOCAP dataset.

语音识别情绪识别自监督学习元数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。