arXiv:2601.15668cs.SD2026-01被引 17

让语音情绪模型像人一样推理,生成可解释的判断。

EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning

  • 用强化学习把情绪识别变成深度推理任务,结合语音韵律线索
  • 在35K条数据上训练,准确率超越现有模型,解释质量显著提升
  • 适合需要可解释性的人机交互、情感计算研究者

语音中的情感信息在多模态感知中具有独特作用。然而,当前语音大模型(SpeechLLMs)与传统语音情绪识别(SER)系统一样,仍将情绪理解视为简单分类问题,导致预测缺乏可解释性,且未发挥大模型的表达与推理能力。本文首次将SER重构为深度推理任务,提出EmotionThinker,通过强化学习生成基于细粒度声学线索的准确情绪预测与可解释解释。我们首先构建了包含思维链标注和详细描述的EmotionCoT-35K情感推理数据集;其次发现现有SpeechLLMs对韵律感知较弱,而韵律是情绪解读的关键信号,因此开发了增强韵律感知的EmotionThinker-Base基础模型;最后提出一种新型强化学习算法GRPO-PTR,其逐步引入推理奖励,根据推理与结果的一致性动态调整信任权重,并基于多维标准评估整体推理质量。EmotionThinker在情绪准确率与解释质量上均优于先前最优模型,推动SER向可解释的多模态推理发展。

原文摘要 · Abstract (English)

Emotional information in speech plays a unique role in multimodal perception. However, current Speech Large Language Models (SpeechLLMs), similar to conventional speech emotion recognition (SER) systems, still treat emotion understanding as a simple classification problem. This provides limited interpretability of predictions, while leaving the LLMs' expressive and reasoning capabilities underutilized. In this work, we take the first step to reformulate SER as a deep reasoning problem through reinforcement learning (RL). We propose EmotionThinker, which is designed to generate accurate emotion predictions with interpretable explanations grounded in fine-grained acoustic cues. To achieve this, we first construct EmotionCoT-35K, an emotional reasoning dataset with Chain-of-Thought annotations and detailed captions. Second, we observe that current SpeechLLMs exhibit weak prosody perception, whereas prosodic cues constitute fundamental signals for interpreting emotions. To address this, we develop the prosody-enhanced foundation model EmotionThinker-Base, and demonstrate that prosody enhancement improves emotion understanding. Third, we introduce Group-Relative-Policy-Optimization with Progressive-Trust-aware-Reasoning-Reward (GRPO-PTR) for RL. Different from standard GRPO, which relies only on rule-based outcome rewards, GRPO-PTR progressively introduces reasoning reward, dynamically adjusts it with a trustworthiness weight reflecting the alignment between reasoning and outcome, and evaluates the overall reasoning quality with a reward model based on multi-dimensional criteria. EmotionThinker outperforms previous state-of-the-art evaluation models both in emotion accuracy and explanation quality, advancing SER toward interpretable multimodal reasoning. Project page: https://github.com/dingdongwang/EmotionThinker

语音情绪可解释性强化学习韵律感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。