基于编码器语言模型的零样本语音克隆系统,实现自然流畅的口语化语音生成。
The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024
- 用带延迟模式的LLaMA编码器语言模型实现零样本口语风格克隆。
- 在约束赛道中达到3.80分的语音自然度评分,表现最优。
- 适合语音克隆、个性化语音合成等场景,尤其关注自然口语表达。
本文介绍了针对ISCSLP 2024对话式语音克隆挑战(CoVoC)的零样本口语风格语音合成系统。我们提出一种基于LLaMA的编码器语言模型,并引入延迟模式以实现自发性语音风格克隆。为提升语音可懂度,我们在语言模型中采用无分类器引导(CFG)策略,增强对标记预测的条件引导。为生成高质量语句,我们实施有效的数据预处理并使用精选高质量口语语音数据进行微调。在CoVoC约束赛道的官方评估中,该系统在语音自然度方面取得3.80分的最高均值意见评分(MOS),并在语音质量和说话人相似性方面获得良好结果。
原文摘要 · Abstract (English)
This paper describes the zero-shot spontaneous style TTS system for the ISCSLP 2024 Conversational Voice Clone Challenge (CoVoC). We propose a LLaMA-based codec language model with a delay pattern to achieve spontaneous style voice cloning. To improve speech intelligibility, we introduce the Classifier-Free Guidance (CFG) strategy in the language model to strengthen conditional guidance on token prediction. To generate high-quality utterances, we adopt effective data preprocessing operations and fine-tune our model with selected high-quality spontaneous speech data. The official evaluations in the CoVoC constrained track show that our system achieves the best speech naturalness MOS of 3.80 and obtains considerable speech quality and speaker similarity results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。