arXiv:2511.08723eess.ASeess.SP2025-11被引 25

提出新框架,让语音对话模型更懂语气和情感。

ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction

  • 用强化学习直接在波形层面优化语音内容与语调。
  • 相比监督微调,响应更贴合语气,提升10%适切性。
  • 适合研究语音交互、情感计算的学者与工程师。

语音到语音(S2S)模型虽展现良好对话能力,但对情绪、语调等副语言线索的处理仍不足,且缺乏高质量表达数据。为此,我们提出基于强化学习的帕拉林语感知S2S框架ParaS2S,直接在波形层面评估并优化响应内容与说话风格。构建了ParaS2SBench基准,采用富有表现力且具挑战性的查询,评估输入输出对在内容与风格上的自然度。设计多阶段多音调训练策略,防止端到端音频大模型判断中的风格幻觉。该自动判别器与人类偏好高度相关且可扩展,使模型能通过强化学习从无标注语音中交互学习。实验表明,现有S2S模型无法恰当响应副语言属性,性能不优于流水线基线。我们的强化学习方法(ParaS2SAlign)在ParaS2SBench上相较监督微调(SFT)实现10%相对改进,超越所有已有模型,且所需成对演示数据显著少于纯SFT。

原文摘要 · Abstract (English)

Speech-to-Speech (S2S) models have shown promising dialogue capabilities, but their ability to handle paralinguistic cues - such as emotion, tone, and speaker attributes - and to respond appropriately in both content and style remains under-explored. Progress is further hindered by the scarcity of high-quality and expressive demonstrations. To address this, we introduce a new reinforcement learning (RL) framework for paralinguistic-aware S2S, ParaS2S, which evaluates and optimizes both response content and speaking style directly at the waveform level. We first construct ParaS2SBench, a benchmark that evaluates the naturalness of input-output pairs in terms of content and speaking style using expressive and challenging queries. For the automatic judge, we propose a PolyTone training strategy and a multi-stage framework, preventing the style hallucination of end-to-end audio LLM judging. Our judge correlates well with human preferences and is scalable, enabling the model to interact and learn from unlabeled speech via RL. Experiments show that existing S2S models fail to respond appropriately to paralinguistic attributes, performing no better than pipeline-based baselines. Our RL approach (ParaS2SAlign) achieves a 10% relative improvement in the appropriateness of response content and speaking style on ParaS2SBench over supervised fine-tuning (SFT), surpassing all prior models while requiring substantially fewer paired demonstrations than pure SFT. Our findings highlight the need for a scalable and accurate automatic evaluator for speech-to-speech interaction.

语音生成强化学习副语言对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。