用偏好对齐和无分类器引导提升大模型语音生成的可控性
Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance
- 结合ASR与声纹验证模型进行偏好对齐,增强生成可控性
- 在小数据集上仍超越现有TTS模型,提升说话人相似度与自然度
- 适合关注语音合成可控性与质量的研究者与开发者
自回归语音标记生成模型虽能生成多样且自然的语音,但缺乏可控性,常出现幻觉或不符合输入条件的发声。我们提出Koel-TTS,一种基于编码器-解码器Transformer的语音合成系统,通过自动语音识别(ASR)和说话人验证模型引导的偏好对齐技术解决此问题。同时引入无分类器引导,进一步提升语音对齐文本和参考音频的程度。实验表明,这些优化显著提升了目标说话人相似度、可懂性和自然度。值得注意的是,Koel-TTS直接将文本和上下文音频映射为声学标记,在上述指标上优于当前最优的TTS模型,且训练数据量显著更小。音频样本与演示可在官网获取。
原文摘要 · Abstract (English)
While autoregressive speech token generation models produce speech with remarkable variety and naturalness, their inherent lack of controllability often results in issues such as hallucinations and undesired vocalizations that do not conform to conditioning inputs. We introduce Koel-TTS, a suite of enhanced encoder-decoder Transformer TTS models that address these challenges by incorporating preference alignment techniques guided by automatic speech recognition and speaker verification models. Additionally, we incorporate classifier-free guidance to further improve synthesis adherence to the transcript and reference speaker audio. Our experiments demonstrate that these optimizations significantly enhance target speaker similarity, intelligibility, and naturalness of synthesized speech. Notably, Koel-TTS directly maps text and context audio to acoustic tokens, and on the aforementioned metrics, outperforms state-of-the-art TTS models, despite being trained on a significantly smaller dataset. Audio samples and demos are available on our website.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。