arXiv:2608.03054cs.SD2026-08

提升语音大模型情感表达力,设计新评测基准与解耦优化方法

Towards More Expressive Spoken LLMs: Fine-Grained Intent Benchmarking and Acoustic-Lexical Decoupled Policy Optimization

论文配图:Towards More Expressive Spoken LLMs: Fine-Grained Intent Benchmarking and Acoustic-Lexical Decoupled Policy Optimization
图 1 · 摘自论文原文
  • 构建中文细粒度情感意图数据集ParaIntent,区分显性与隐性表达
  • 提出声学-词汇解耦策略,使语音与文本分别优化,情感表达更自然
  • 在合成与真人录音测试中均显著提升情感表现,适合语音对话系统研究者

语音情感对话要求模型理解用户语音输入并生成语义恰当且情感丰富的回应。难点在于意图可能通过词汇明确表达,或通过语气等副语言特征隐含传达,二者可能一致或冲突。当前主要受限于缺乏区分此类意图表达的评测基准,以及未能同时兼顾响应质量与情感表达的强化学习目标。为此,我们提出ParaIntent,一个包含14类意图、显性与隐性样本均衡的中文基准,配备多维度评估协议,涵盖意图达成、响应质量与情感表达。针对策略优化,现有方法或共享文本与语音目标,或仅对单一模态使用强化学习,导致模态信号在策略中纠缠。我们提出声学-词汇解耦策略(ALPO),独立计算文本与声学优势,并分别作用于对应文本与语音标记,在统一采样框架下实现。在相同奖励函数与训练预算下,ALPO在多数自动指标上优于标准GRPO,主观评价最优,尤其在合成与真人录音测试集中情感表达提升显著。

原文摘要 · Abstract (English)

Spoken emotional dialogue requires a model to understand a user's spoken input and generate a response that is both semantically appropriate and emotionally expressive. This is challenging because communicative intent may be stated explicitly in lexical content or conveyed more implicitly through paralinguistic cues, which can complement or diverge from the words themselves. However, two limitations constrain progress in this area: the scarcity of benchmarks that distinguish these intent expressions, and the lack of reinforcement learning objectives that jointly account for response quality and emotional expression. To address the lack of suitable benchmarks, we introduce ParaIntent, a Chinese benchmark comprising 14 intent categories with balanced explicit and implicit samples, together with a multidimensional evaluation protocol covering intent fulfillment, response quality, and emotional expression. For policy optimization, existing approaches either use a shared objective for text and speech or apply reinforcement learning to only one modality, leaving modality-specific learning signals entangled within policy optimization. Motivated by this, we propose Acoustic-Lexical Decoupled Policy Optimization (ALPO), which computes independent textual and acoustic advantages and routes them to the corresponding text and speech tokens within a unified rollout. Under identical reward functions and training budgets, ALPO improves over standard GRPO on most automatic metrics and achieves the best subjective results among the fine-tuned variants, with particularly clear gains in emotional expressiveness on both the synthetic and human-recorded test sets.

语音生成情感对话强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。