用户手绘语调草图,即可精准控制语音情感与细节。
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
- 用用户手绘的语调草图作为控制信号,生成语音。
- 能精确还原音高和能量的细微变化,实现细粒度控制。
- 适合需要个性化语音表达的创作者与开发者。
如何让文本转语音系统按用户预期的语调合成语音,是当前研究热点。现有方法主要依赖参考语音或自然语言描述,但前者需大量筛选,后者仅能控制整体语调,难以实现精细调控。本文提出DrawSpeech,一种基于语调草图的扩散模型,用户可绘制粗略的语调趋势图,模型据此恢复详细的音高与能量轮廓,并生成目标语音。实验表明,DrawSpeech能生成多样化的语调,支持用户友好的细粒度控制。代码与音频样本已公开。
原文摘要 · Abstract (English)
Controlling text-to-speech (TTS) systems to synthesize speech with the prosodic characteristics expected by users has attracted much attention. To achieve controllability, current studies focus on two main directions: (1) using reference speech as prosody prompt to guide speech synthesis, and (2) using natural language descriptions to control the generation process. However, finding reference speech that exactly contains the prosody that users want to synthesize takes a lot of effort. Description-based guidance in TTS systems can only determine the overall prosody, which has difficulty in achieving fine-grained prosody control over the synthesized speech. In this paper, we propose DrawSpeech, a sketch-conditioned diffusion model capable of generating speech based on any prosody sketches drawn by users. Specifically, the prosody sketches are fed to DrawSpeech to provide a rough indication of the expected prosody trends. DrawSpeech then recovers the detailed pitch and energy contours based on the coarse sketches and synthesizes the desired speech. Experimental results show that DrawSpeech can generate speech with a wide variety of prosody and can precisely control the fine-grained prosody in a user-friendly manner. Our implementation and audio samples are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。