arXiv:2505.23236cs.SDcs.HC2025-05中稿 · INTERSPEECH2025被引 5

用大模型拆解语音情绪细节,让识别结果更可信。

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition

  • 用大模型交替优化情绪识别与语音特征分离
  • 在IEMOCAP和MELD上准确率提升超3.7%以上
  • 可解释的语音特征适合需要透明性的场景

本文提出一种端到端的、由大语言模型驱动的可解释语音情绪识别方法。通过交替微调大模型,从HuBERT自监督表示中解耦出精细的语音情绪描述符(SED)特征,如音高、语调和重音,并联合进行情绪识别与语音识别任务。利用信息瓶颈(IB)压缩的变分自编码器(VAE)提取的HuBERT特征,调节特征粒度。在IEMOCAP和MELD基准上的实验表明,该方法持续优于基于LLaMA的对比基线,包括仅使用交替多任务微调或仅特征解耦的方法。情绪识别的未加权准确率绝对提升最高达4.0%和3.7%,相对提升5.4%和6.6%。更重要的是,情绪描述符提供了额外的可解释性支持。

原文摘要 · Abstract (English)

This paper presents a novel end-to-end LLM-empowered explainable speech emotion recognition (SER) approach. Fine-grained speech emotion descriptor (SED) features, e.g., pitch, tone and emphasis, are disentangled from HuBERT SSL representations via alternating LLM fine-tuning to joint SER-SED prediction and ASR tasks. VAE compressed HuBERT features derived via Information Bottleneck (IB) are used to adjust feature granularity. Experiments on the IEMOCAP and MELD benchmarks demonstrate that our approach consistently outperforms comparable LLaMA-based SER baselines, including those using either (a) alternating multi-task fine-tuning alone or (b) feature disentanglement only. Statistically significant increase of SER unweighted accuracy by up to 4.0% and 3.7% absolute (5.4% and 6.6% relative) are obtained. More importantly, emotion descriptors offer further explainability for SER.

情绪识别大模型可解释性语音分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。