arXiv:2510.14628cs.CLcs.AI2025-10被引 1

用AI反馈让语音合成更富情感,无需人工标注。

RLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech Synthesis

  • 用AI自动评估语义和韵律情感一致性,替代人工标注。
  • 在多个数据集上降低26.1%错词率,主观评分提升超10%。
  • 适合需要高情感表达的语音合成场景,如对话系统、有声书。

近年来,文本转语音(TTS)技术在中性语调下已达到接近人类水平的语音质量。然而,现有方法大多依赖昂贵的情感标注,或优化无法准确反映感知情感质量的代理目标,导致生成语音虽语义正确,却缺乏表现力和情感丰富性。为此,我们提出 RLAIF-SPA 框架,通过强化学习从 AI 反馈中直接优化情感表现力与可懂性,无需人工监督。该框架结合自动语音识别(ASR)提供语义准确性反馈,并采用结构化奖励建模评估韵律-情感一致性。RVAIF-SPA 实现了对表达性语音生成在结构、情感、语速和语调四个维度的精细控制。在 Libri-Speech、MELD 及中文 ESD 数据集上的大量实验表明,该方法在朗读、对话及情感语音任务中均取得稳定提升。在 Libri-Speech 上,其性能持续优于 Chat-TTS,词错误率降低 26.1%,SIM-O 提升 9.1%,人主观评价提升超过 10%。

原文摘要 · Abstract (English)

Recent advances in Text-To-Speech (TTS) synthesis have achieved near-human speech quality in neutral speaking styles. However, most existing approaches either depend on costly emotion annotations or optimize surrogate objectives that fail to adequately capture perceptual emotional quality. As a result, the generated speech, while semantically accurate, often lacks expressive and emotionally rich characteristics. To address these limitations, we propose RLAIF-SPA, a novel framework that integrates Reinforcement Learning from AI Feedback (RLAIF) to directly optimize both emotional expressiveness and intelligibility without human supervision. Specifically, RLAIF-SPA incorporates Automatic Speech Recognition (ASR) to provide semantic accuracy feedback, while leveraging structured reward modeling to evaluate prosodic-emotional consistency. RLAIF-SPA enables more precise and nuanced control over expressive speech generation along four structured evaluation dimensions: Structure, Emotion, Speed, and Tone. Extensive experiments on Libri-Speech, MELD, and Mandarin ESD datasets demonstrate consistent gains across clean read speech, conversational dialogue, and emotional speech. On Libri-Speech, RLAIF-SPA consistently outperforms Chat-TTS, achieving a 26.1% reduction in word error rate, a 9.1% improvement in SIM-O, and over 10% gains in human subjective evaluations.

语音合成情感表达强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。