用结构化情感图提升语音情绪识别的零样本表现
Plug-and-Play Emotion Graphs for Compositional Prompting in Zero-Shot Speech Emotion Recognition
- 引入情感图结构,融合音调、语速等七类声学特征
- 在多个基准上超越现有最优方法,显著提升识别准确率
- 无需微调,适合希望快速部署的语音分析应用
大型音频-语言模型(LALMs)在语音任务中展现出强大的零样本能力,但在语音情绪识别(SER)任务中因声调建模弱和跨模态推理有限而表现不佳。我们提出一种组合式思维链情感推理框架(CCoT-Emo),通过引入结构化情感图(EGs)引导LALMs进行情绪推断,无需微调。每个情感图编码七种声学特征(如音调、语速、抖动、闪烁)、文本情感、关键词及跨模态关联关系。将情感图嵌入提示词后,提供可解释且组合化的表示,增强LALM的推理能力。在多个SER基准上的实验表明,CCoT-Emo优于先前最先进方法,并显著提升零样本基线的准确率。
原文摘要 · Abstract (English)
Large audio-language models (LALMs) exhibit strong zero-shot performance across speech tasks but struggle with speech emotion recognition (SER) due to weak paralinguistic modeling and limited cross-modal reasoning. We propose Compositional Chain-of-Thought Prompting for Emotion Reasoning (CCoT-Emo), a framework that introduces structured Emotion Graphs (EGs) to guide LALMs in emotion inference without fine-tuning. Each EG encodes seven acoustic features (e.g., pitch, speech rate, jitter, shimmer), textual sentiment, keywords, and cross-modal associations. Embedded into prompts, EGs provide interpretable and compositional representations that enhance LALM reasoning. Experiments across SER benchmarks show that CCoT-Emo outperforms prior SOTA and improves accuracy over zero-shot baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。