arXiv:2510.10078cs.SDcs.LG2025-10被引 1

用互信息正则化生成模型,提升语音情感识别的数据质量与性能

Improving Speech Emotion Recognition with Mutual Information Regularized Generative Model

  • 基于互信息正则化生成情感特征,确保生成内容与标签强关联
  • 在三个数据集上实现最高2.6%的单模态和3.2%的多模态性能提升
  • 适用于情感计算中数据稀缺场景,尤其适合复杂模型训练

大规模、高质量的情感语音语料库的缺乏持续限制了语音情感识别(SER)的性能与鲁棒性,尤其在模型复杂度上升和多模态系统需求增加的背景下。生成式数据增强提供了一种有前景的解决方案,但现有方法常因对类别标签的简化条件处理而生成情感不一致的样本。本文提出一种新型互信息正则化生成框架,结合跨模态对齐与特征级合成。基于InfoGAN架构,首先利用预训练变压器和对比目标学习语义对齐的音视频表示空间。随后训练特征生成器,生成具有情感感知的音频特征,并以互信息作为定量正则项,确保生成特征与其条件变量间强依赖关系。该方法拓展至多模态场景,可生成新颖的(音频,文本)配对特征。在IEMOCAP、MSP-IMPROV和MSP-Podcast三个基准数据集上的综合评估表明,本框架始终优于现有增强方法,在单模态SER中性能提升达2.6%,多模态情感识别提升达3.2%。更重要的是,我们证明互信息兼具正则化与生成质量度量功能,为情感计算中的数据增强提供了系统性方法。

原文摘要 · Abstract (English)

Lack of large, well-annotated emotional speech corpora continues to limit the performance and robustness of speech emotion recognition (SER), particularly as models grow more complex and the demand for multimodal systems increases. While generative data augmentation offers a promising solution, existing approaches often produce emotionally inconsistent samples due to oversimplified conditioning on categorical labels. This paper introduces a novel mutual-information-regularised generative framework that combines cross-modal alignment with feature-level synthesis. Building on an InfoGAN-style architecture, our method first learns a semantically aligned audio-text representation space using pre-trained transformers and contrastive objectives. A feature generator is then trained to produce emotion-aware audio features while employing mutual information as a quantitative regulariser to ensure strong dependency between generated features and their conditioning variables. We extend this approach to multimodal settings, enabling the generation of novel, paired (audio, text) features. Comprehensive evaluation on three benchmark datasets (IEMOCAP, MSP-IMPROV, MSP-Podcast) demonstrates that our framework consistently outperforms existing augmentation methods, achieving state-of-the-art performance with improvements of up to 2.6% in unimodal SER and 3.2% in multimodal emotion recognition. Most importantly, we demonstrate that mutual information functions as both a regulariser and a measurable metric for generative quality, offering a systematic approach to data augmentation in affective computing.

语音情感识别生成模型互信息数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。