arXiv:2602.07803eess.AScs.AI2026-02被引 6

开源高质量歌声合成系统,支持零样本跨语言生成。

SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis

  • 基于4.2万小时数据训练,支持中英粤三语零样本合成
  • 可按乐谱或旋律控制,合成质量达当前最优水平
  • 配套专用评估集,支持真实场景下的零样本测试

近年来语音合成技术快速发展,但开源歌声合成(SVS)系统在实际部署中仍面临鲁棒性和零样本泛化能力不足的挑战。本文提出SoulX-Singer,一个面向实际应用的高质量开源歌声合成系统。该系统支持基于符号音乐乐谱(MIDI)或旋律表示的可控歌声生成,适用于真实生产工作流。模型在超过42,000小时的声乐数据上训练,支持中文、英文和粤语,在多种音乐条件下均保持领先合成质量。此外,为实现零样本歌声合成在实际场景中的可靠评估,我们构建了SoulX-Singer-Eval基准,采用严格的训练-测试分离设计,支持系统性地评估零样本设置下的性能。

原文摘要 · Abstract (English)

While recent years have witnessed rapid progress in speech synthesis, open-source singing voice synthesis (SVS) systems still face significant barriers to industrial deployment, particularly in terms of robustness and zero-shot generalization. In this report, we introduce SoulX-Singer, a high-quality open-source SVS system designed with practical deployment considerations in mind. SoulX-Singer supports controllable singing generation conditioned on either symbolic musical scores (MIDI) or melodic representations, enabling flexible and expressive control in real-world production workflows. Trained on more than 42,000 hours of vocal data, the system supports Mandarin Chinese, English, and Cantonese and consistently achieves state-of-the-art synthesis quality across languages under diverse musical conditions. Furthermore, to enable reliable evaluation of zero-shot SVS performance in practical scenarios, we construct SoulX-Singer-Eval, a dedicated benchmark with strict training-test disentanglement, facilitating systematic assessment in zero-shot settings.

歌声合成零样本多语言开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。