ConSinger用最少步数实现高质量歌声合成,兼顾速度与音质。
ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps
- 基于一致性模型,通过约束训练提升生成效率。
- 仅需少量推理步骤即可生成高保真歌声,速度优于传统扩散模型。
- 适合对实时性要求高的语音合成应用,如音乐生成与交互系统。
歌唱语音合成(SVS)系统旨在根据乐谱(歌词、时长和音高)生成高保真歌声。近年来,扩散模型在此领域表现优异,但为换取高质量输出而牺牲推理速度,限制了其应用场景。为更高效地生成高质量歌声,本文提出基于一致性模型的歌唱语音合成方法ConSinger,实现最小步数下的高保真合成。模型通过施加一致性约束进行训练,生成质量显著提升,仅以少量速度损失为代价。实验表明,ConSinger在生成速度和质量上均具备强竞争力。音频样例可访问 https://keylxiao.github.io/consinger。
原文摘要 · Abstract (English)
Singing voice synthesis (SVS) system is expected to generate high-fidelity singing voice from given music scores (lyrics, duration and pitch). Recently, diffusion models have performed well in this field. However, sacrificing inference speed to exchange with high-quality sample generation limits its application scenarios. In order to obtain high quality synthetic singing voice more efficiently, we propose a singing voice synthesis method based on the consistency model, ConSinger, to achieve high-fidelity singing voice synthesis with minimal steps. The model is trained by applying consistency constraint and the generation quality is greatly improved at the expense of a small amount of inference speed. Our experiments show that ConSinger is highly competitive with the baseline model in terms of generation speed and quality. Audio samples are available at https://keylxiao.github.io/consinger.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。