arXiv:2502.18186cs.SDcs.CL2025-02被引 13

用上下文感知和思维链提升大模型语音情感识别稳定性

Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought

  • 融合语义与声学特征,分步推理增强情感识别
  • 在IEMOCAP数据集上准确率达78.6%,优于Qwen2-Audio等模型
  • 适合语音情感分析、人机交互领域研究者参考

大规模音频语言模型(ALM)如Qwen2-Audio能理解多种音频信号,实现音频分析并生成文本响应。但在语音情感识别(SER)任务中,这类模型常因幻觉导致误判或无关输出。为此,本文提出C²SER,一种通过上下文感知与思维链(CoT)提升SER稳定性和准确性的新ALM。C²SER整合Whisper编码器进行语义感知,以及基于半监督学习扩展的Emotion2Vec-S实现声学感知,以增强情感区分能力。同时采用分步思维链推理,结合语音内容与表达风格优化识别。为提升稳定性,引入显式到隐式思维链的自蒸馏机制,减少误差累积。大量实验表明,C²SER在IEMOCAP数据集上达到78.6%的准确率,优于Qwen2-Audio和SECap等主流模型。论文开源了训练代码、模型权重及测试集,促进后续研究。

原文摘要 · Abstract (English)

Large-scale audio language models (ALMs), such as Qwen2-Audio, are capable of comprehending diverse audio signal, performing audio analysis and generating textual responses. However, in speech emotion recognition (SER), ALMs often suffer from hallucinations, resulting in misclassifications or irrelevant outputs. To address these challenges, we propose C$^2$SER, a novel ALM designed to enhance the stability and accuracy of SER through Contextual perception and Chain of Thought (CoT). C$^2$SER integrates the Whisper encoder for semantic perception and Emotion2Vec-S for acoustic perception, where Emotion2Vec-S extends Emotion2Vec with semi-supervised learning to enhance emotional discrimination. Additionally, C$^2$SER employs a CoT approach, processing SER in a step-by-step manner while leveraging speech content and speaking styles to improve recognition. To further enhance stability, C$^2$SER introduces self-distillation from explicit CoT to implicit CoT, mitigating error accumulation and boosting recognition accuracy. Extensive experiments show that C$^2$SER outperforms existing popular ALMs, such as Qwen2-Audio and SECap, delivering more stable and precise emotion recognition. We release the training code, checkpoints, and test sets to facilitate further research.

语音情感识别大模型思维链音频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。