arXiv:2509.15654cs.SDeess.AS2025-09EMNLP被引 8

用情绪规则强化学习,提升语音情感识别的推理与泛化能力

EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognition

  • 引入情绪相似性奖励和结构化推理机制
  • 在MELD和IEMOCAP上达顶尖性能,跨数据集泛化强
  • 适合需要精准情感分析与可解释推理的场景

尽管大型音频-语言模型(LALMs)在听觉理解任务中表现优异,但在情感计算场景下,其情感识别、推理及细微情绪区分能力仍不理想。近期强化学习(RL)进展显示其可提升LALMs的推理能力,但直接应用于语音情感识别(SER)面临两大挑战:(1)情绪边界模糊导致收敛不稳定;(2)小模型(如7B参数架构)推理能力有限。为此,我们提出EMO-RL框架,融合强化学习并引入两项创新:情绪相似性加权奖励(ESWR)与显式结构化推理(ESR)。基于预训练的LALMs,采用群体相对策略优化并施加情绪约束。大量实验表明,该方法显著增强LALMs的情感推理能力,在MELD与IEMOCAP数据集上均达到当前最优表现,跨数据集实验进一步验证其强大泛化性。

原文摘要 · Abstract (English)

Although Large Audio-Language Models (LALMs) have exhibited outstanding performance in auditory understanding, their performance in affective computing scenarios, particularly in emotion recognition, reasoning, and subtle sentiment differentiation, remains suboptimal. Recent advances in Reinforcement Learning (RL) have shown promise in improving LALMs' reasoning abilities. However, two critical challenges hinder the direct application of RL techniques to Speech Emotion Recognition (SER) tasks: (1) convergence instability caused by ambiguous emotional boundaries and (2) limited reasoning ability when using relatively small models (e.g., 7B-parameter architectures). To overcome these limitations, we introduce EMO-RL, a novel framework incorporating reinforcement learning with two key innovations: Emotion Similarity-Weighted Reward (ESWR) and Explicit Structured Reasoning (ESR). Built upon pretrained LALMs, our method employs group-relative policy optimization with emotion constraints. Comprehensive experiments demonstrate that our EMO-RL training strategies can significantly enhance the emotional reasoning capabilities of LALMs, attaining state-of-the-art results on both the MELD and IEMOCAP datasets, and cross-dataset experiments prove the strong superiority of generalization.

语音情感识别强化学习大模型推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。