arXiv:2410.20336cs.CLcs.AI2024-10被引 7

用大模型生成自然语音,实现文本到语音的高质量合成。

Get Large Language Models Ready to Speak: A Late-fusion Approach for Speech Generation

  • 通过微调Llama模型构建文本转语音系统,支持语音生成。
  • 提出多专家架构的双模态模型,兼顾文本问答与语音合成性能。
  • 采用后期融合方式避免遗忘,适合需要语音交互的场景。

大型语言模型(LLM)在各类文本任务中表现出色,但将其拓展至语音生成仍鲜有探索。本文提出基于微调Llama模型的文本转语音系统TTS-Llama,实现业界领先语音合成效果。在此基础上,进一步构建了通过纯后期融合参数高效微调(PEFT)与多专家架构训练的文本-语音多模态大模型MoLE-Llama。大量实验表明,MoLE-Llama在仅文本问答(QA)和语音合成任务上均表现优异,有效缓解单一模态遗忘问题。最后,我们探索其在‘文本输入、语音输出’问答任务中的应用,验证其作为具备语音生成能力的多模态对话系统的巨大潜力。

原文摘要 · Abstract (English)

Large language models (LLMs) have revolutionized natural language processing (NLP) with impressive performance across various text-based tasks. However, the extension of text-dominant LLMs to with speech generation tasks remains under-explored. In this work, we introduce a text-to-speech (TTS) system powered by a fine-tuned Llama model, named TTS-Llama, that achieves state-of-the-art speech synthesis performance. Building on TTS-Llama, we further propose MoLE-Llama, a text-and-speech multimodal LLM developed through purely late-fusion parameter-efficient fine-tuning (PEFT) and a mixture-of-expert architecture. Extensive empirical results demonstrate MoLE-Llama's competitive performance on both text-only question-answering (QA) and TTS tasks, mitigating catastrophic forgetting issue in either modality. Finally, we further explore MoLE-Llama in text-in-speech-out QA tasks, demonstrating its great potential as a multimodal dialog system capable of speech generation.

语音生成多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。