arXiv:2409.03283cs.SDeess.AS2024-09被引 124

FireRedTTS打造工业级语音生成框架,支持零样本配音与类人对话机器人。

FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

  • 用语言模型驱动语音生成,先分词再合成高保真波形。
  • 零样本克隆用户语音,1小时微调即可适配专业配音角色。
  • 支持带情绪和语调的自然对话生成,适合聊天机器人场景。

本文提出FireRedTTS,一个面向产业级生成式语音应用的基座文本转语音框架。该框架包含数据处理、基座系统与下游应用三部分。首先,构建了大规模高质量语音数据集,涵盖丰富内容、语调与音色。其次,提出基于语言模型的基座TTS系统:通过语义感知语音分词器将声波压缩为离散语义令牌,由语言模型根据文本与音频提示生成,并经两阶段波形生成器还原为高保真语音。最后,展示两个应用场景:零样本语音克隆用于UGC配音,少量微调(1小时录音)即可适配专业表达风格;通过指令微调实现可控类人语音生成,支持自然语气与情感表达,适用于对话机器人。

原文摘要 · Abstract (English)

This work proposes FireRedTTS, a foundation text-to-speech framework, to meet the growing demands for personalized and diverse generative speech applications. The framework comprises three parts: data processing, foundation system, and downstream applications. First, we comprehensively present our data processing pipeline, which transforms massive raw audio into a large-scale high-quality TTS dataset with rich annotations and a wide coverage of content, speaking style, and timbre. Then, we propose a language-model-based foundation TTS system. The speech signal is compressed into discrete semantic tokens via a semantic-aware speech tokenizer, and can be generated by a language model from the prompt text and audio. Then, a two-stage waveform generator is proposed to decode them to the high-fidelity waveform. We present two applications of this system: voice cloning for dubbing and human-like speech generation for chatbots. The experimental results demonstrate the solid in-context learning capability of FireRedTTS, which can stably synthesize high-quality speech consistent with the prompt text and audio. For dubbing, FireRedTTS can clone target voices in a zero-shot way for the UGC scenario and adapt to studio-level expressive voice characters in the PUGC scenario via few-shot fine-tuning with 1-hour recording. Moreover, FireRedTTS achieves controllable human-like speech generation in a casual style with paralinguistic behaviors and emotions via instruction tuning, to better serve spoken chatbots.

语音生成文本转语音语音克隆对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。