首个基于变分自编码器的端到端流式歌声合成系统,低延迟且音色自然。
CSSinger: End-to-End Chunkwise Streaming Singing Voice Synthesis System Based on Conditional Variational Autoencoder
- 采用条件变分自编码器,直接从乐谱生成音频,跳过中间特征步骤。
- 在流式推理下仍保持高表现力和音高准确性,适合实时应用。
- 首次实现基于隐变量的端到端流式音频合成,适合音乐生成与语音合成研究者。
歌声合成(SVS)旨在生成高保真且富有表现力的歌声。传统方法通常先通过声学模型将乐谱转换为声学特征,再由声码器重建歌声。近年来,端到端建模在歌声合成与文本转语音(TTS)领域表现出色。本文提出一种完全端到端的歌声合成方法,并引入分块流式推理以解决实际应用中的延迟问题。这是首次在变分自编码器(VAE)中实现基于隐表示的完整端到端流式音频合成。我们针对流式歌声合成中的隐表示性能进行了专门优化。实验结果表明,该方法在流式歌声合成与文本转语音任务中均实现了高表现力和精确的音高控制。
原文摘要 · Abstract (English)
Singing Voice Synthesis (SVS) aims to generate singing voices of high fidelity and expressiveness. Conventional SVS systems usually utilize an acoustic model to transform a music score into acoustic features, followed by a vocoder to reconstruct the singing voice. It was recently shown that end-to-end modeling is effective in the fields of SVS and Text to Speech (TTS). In this work, we thus present a fully end-to-end SVS method together with a chunkwise streaming inference to address the latency issue for practical usages. Note that this is the first attempt to fully implement end-to-end streaming audio synthesis using latent representations in VAE. We have made specific improvements to enhance the performance of streaming SVS using latent representations. Experimental results demonstrate that the proposed method achieves synthesized audio with high expressiveness and pitch accuracy in both streaming SVS and TTS tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。