arXiv:2512.14657cs.SD2025-12中稿 · NeurIPS

用1.7B语音模型实现歌声合成,仅需135小时数据。

Adapting Speech Language Model to Singing Voice Synthesis

论文配图:Adapting Speech Language Model to Singing Voice Synthesis
图 1 · 摘自论文原文
  • 用音乐谱和歌声构建多流语言模型,预测音高与节奏。
  • 在ACE-Opencpop数据上达到主流离散编码模型水平。
  • 适合想用通用语音模型做歌声生成的研究者。

语音语言模型(SLMs)近年来成为处理多种语音任务的统一范式,包括文本到语音(TTS)、语音增强(SE)和自动语音识别(ASR)。然而,大规模预训练的SLMs泛化能力仍待探索。本文将一个1.7B参数的预训练TTS SLM用于歌唱语音合成(SVS),仅使用135小时的合成歌唱语料库ACE-Opencpop。基于ESPNet-SpeechLM,方法包含:(1)对乐谱条件和歌唱波形进行分词;(2)多流语言模型预测音素;(3)基于条件流匹配的梅尔频谱生成;(4)梅尔到波形的声码器。实验表明,该适配后的模型在SVS任务上表现良好,性能可媲美领先的离散令牌型SVS模型。

原文摘要 · Abstract (English)

Speech Language Models (SLMs) have recently emerged as a unified paradigm for addressing a wide range of speech-related tasks, including text-to-speech (TTS), speech enhancement (SE), and automatic speech recognition (ASR). However, the generalization capability of large-scale pre-trained SLMs remains underexplored. In this work, we adapt a 1.7B parameter TTS pretrained SLM for singing voice synthesis (SVS), using only a 135-hour synthetic singing corpus, ACE-Opencpop. Building upon the ESPNet-SpeechLM, our recipe involves the following procedure: (1) tokenization of music score conditions and singing waveforms, (2) multi-stream language model token prediction, (3) conditional flow matching-based mel-spectrogram generation. (4) a mel-to-wave vocoder. Experimental results demonstrate that our adapted SLM generalizes well to SVS and achieves performance comparable to leading discrete token-based SVS models.

歌声合成语音模型音乐生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。