arXiv:2603.15981cs.CLcs.AI2026-03Conference of the …被引 2

用强化学习让语音大模型更懂语气情绪,提升理解与生成能力。

Aligning Paralinguistic Understanding and Generation in Speech LLMs via Multi-Task Reinforcement Learning

  • 多任务强化学习结合思维链提示,显式建模情感推理过程。
  • 在Expresso、IEMOCAP、RAVDESS上比主流模型提升8-12%准确率。
  • 适合需要高情商语音交互的场景,如客服、心理陪伴。

语音大语言模型能感知语调、情绪和非语言声音等副语言线索,这对意图理解至关重要。然而,利用这些线索面临训练数据少、标注难以及模型依赖词汇捷径等问题。本文提出一种结合思维链提示的多任务强化学习方法,以激发显式的感情推理。为缓解数据稀缺,我们构建了副语言感知的语音大模型(PALLM),通过两阶段流程联合优化音频情感分类与副语言感知的回答生成。实验表明,该方法在Expresso、IEMOCAP和RAVDESS数据集上,相比监督基线及强大多媒体模型(Gemini-2.5-Pro、GPT-4o-audio)提升了8-12%的性能。结果证明,采用多任务强化学习建模副语言推理,对构建具备情感智能的语音大模型至关重要。

原文摘要 · Abstract (English)

Speech large language models (LLMs) observe paralinguistic cues such as prosody, emotion, and non-verbal sounds--crucial for intent understanding. However, leveraging these cues faces challenges: limited training data, annotation difficulty, and models exploiting lexical shortcuts over paralinguistic signals. We propose multi-task reinforcement learning (RL) with chain-of-thought prompting that elicits explicit affective reasoning. To address data scarcity, we introduce a paralinguistics-aware speech LLM (PALLM) that jointly optimizes sentiment classification from audio and paralinguistics-aware response generation via a two-stage pipeline. Experiments demonstrate that our approach improves paralinguistics understanding over both supervised baselines and strong proprietary models (Gemini-2.5-Pro, GPT-4o-audio) by 8-12% on Expresso, IEMOCAP, and RAVDESS. The results show that modeling paralinguistic reasoning with multi-task RL is crucial for building emotionally intelligent speech LLMs.

语音大模型情感理解强化学习副语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。