arXiv:2504.19146cs.SDeess.AS2025-04被引 4

开源低成本语音合成模型,专为播客场景定制,支持快速个性化声音克隆。

Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget

  • 基于超10万小时播客数据预训练,实现零样本高质量语音生成
  • 仅需几十分钟目标语音即可完成个性声音适配,适合个人化播客应用
  • 全链路开源,含训练、推理优化框架,预算控制在5万美元内

近年来,文本到语音(TTS)模型的发展得益于大语言模型(LLM)的融合,提升了语义理解与语音自然度。然而,现有基于LLM的TTS模型普遍缺乏开源训练代码和高效的推理加速框架,限制了其可访问性与可定制性。此外,尚无公开可用的专为播客场景优化的TTS模型,而该类需求在语音交互应用中日益增长。为此,我们提出Muyan-TTS,一个在5万美元预算内、面向播客场景的开源可训练TTS模型。该模型在超过10万小时播客音频数据上预训练,支持零样本语音合成并生成高质量语音。同时,仅需数十分钟的目标语音即可实现说话人自适应,显著提升个性化表达能力。除模型外,我们还开源了完整的数据收集与处理流程、训练方案及优化推理框架,全面支持高效部署。代码与模型已发布于 https://github.com/MYZY-AI/Muyan-TTS。

原文摘要 · Abstract (English)

Recent advancements in text-to-speech (TTS) models have been driven by the integration of large language models (LLMs), enhancing semantic comprehension and improving speech naturalness. However, existing LLM-based TTS models often lack open-source training code and efficient inference acceleration frameworks, limiting their accessibility and adaptability. Additionally, there is no publicly available TTS model specifically optimized for podcast scenarios, which are in high demand for voice interaction applications. To address these limitations, we introduce Muyan-TTS, an open-source trainable TTS model designed for podcast applications within a $50,000 budget. Our model is pre-trained on over 100,000 hours of podcast audio data, enabling zero-shot TTS synthesis with high-quality voice generation. Furthermore, Muyan-TTS supports speaker adaptation with dozens of minutes of target speech, making it highly customizable for individual voices. In addition to open-sourcing the model, we provide a comprehensive data collection and processing pipeline, a full training procedure, and an optimized inference framework that accelerates LLM-based TTS synthesis. Our code and models are available at https://github.com/MYZY-AI/Muyan-TTS.

语音合成播客应用开源模型个性语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。