用扩散模型生成多样且自然的对话语音,提升表达力与上下文一致性。
DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models
- 基于扩散模型预测多样的语调嵌入,融合多模态对话上下文。
- 采用语言模型驱动的语音合成骨架,提升语音自然度与质量。
- 适合需要高表达力和多样性的对话系统研发者使用。
对话语音合成(CSS)旨在生成既符合语境又富有表现力的语音,已有研究致力于增强对对话上下文的理解。然而,现有系统多为确定性预测,忽视了回应的多样性;同时极少采用基于语言模型(LM)的语音合成(TTS)骨干网络,限制了语音的自然度与质量。为此,本文提出DiffCSS,一种创新的CSS框架,结合扩散模型与基于语言模型的TTS骨干网络,生成多样、富有表现力且语境一致的语音。该框架首先设计了一个基于扩散的上下文感知语调预测器,可基于多模态对话上下文采样多样语调嵌入;随后开发一个语调可控的基于语言模型的TTS骨干网络,利用采样的语调嵌入合成高质量语音。实验结果表明,DiffCSS生成的语音在多样性、语境连贯性与表现力上均优于现有系统。
原文摘要 · Abstract (English)
Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding of conversational context. However, existing CSS systems are limited to deterministic prediction, overlooking the diversity of potential responses. Moreover, they rarely employ language model (LM)-based TTS backbones, limiting the naturalness and quality of synthesized speech. To address these issues, in this paper, we propose DiffCSS, an innovative CSS framework that leverages diffusion models and an LM-based TTS backbone to generate diverse, expressive, and contextually coherent speech. A diffusion-based context-aware prosody predictor is proposed to sample diverse prosody embeddings conditioned on multimodal conversational context. Then a prosody-controllable LM-based TTS backbone is developed to synthesize high-quality speech with sampled prosody embeddings. Experimental results demonstrate that the synthesized speech from DiffCSS is more diverse, contextually coherent, and expressive than existing CSS systems
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。