零样本生成高质量播客,支持长时长多角色即兴对话。
MoonCast: High-Quality Zero-Shot Podcast Generation
- 用长上下文语言模型建模音频,突破传统语音合成时长限制。
- 引入播客生成模块,使输出更自然、有即兴感,提升连贯性。
- 无需见过的说话人声音,即可生成逼真播客,适合内容创作者。
近期文本转语音技术在生成高质量短句方面取得显著进展,但在处理长时长、多说话人、即兴对话等真实场景(如播客)时仍面临挑战。主要难点在于:1)长音频:播客通常持续数分钟,超出多数现有方法的上限;2)即兴性:播客具有口语化、非正式特征,与书面语训练数据差异大。本文提出 MoonCast,一种高保真零样本播客生成方案,可将纯文本(如故事、报告、网页)转化为由未见说话人音色驱动的自然播客语音。为生成长音频,采用基于大规模长上下文语音数据的长上下文语言模型音频建模;为增强即兴性,设计播客生成模块,生成包含自然停顿、语气词等细节的脚本,实证表明其重要性不亚于语音合成本身。实验显示,MoonCast 在即兴性和连贯性上显著优于基线。
原文摘要 · Abstract (English)
Recent advances in text-to-speech synthesis have achieved notable success in generating high-quality short utterances for individual speakers. However, these systems still face challenges when extending their capabilities to long, multi-speaker, and spontaneous dialogues, typical of real-world scenarios such as podcasts. These limitations arise from two primary challenges: 1) long speech: podcasts typically span several minutes, exceeding the upper limit of most existing work; 2) spontaneity: podcasts are marked by their spontaneous, oral nature, which sharply contrasts with formal, written contexts; existing works often fall short in capturing this spontaneity. In this paper, we propose MoonCast, a solution for high-quality zero-shot podcast generation, aiming to synthesize natural podcast-style speech from text-only sources (e.g., stories, technical reports, news in TXT, PDF, or Web URL formats) using the voices of unseen speakers. To generate long audio, we adopt a long-context language model-based audio modeling approach utilizing large-scale long-context speech data. To enhance spontaneity, we utilize a podcast generation module to generate scripts with spontaneous details, which have been empirically shown to be as crucial as the text-to-speech modeling itself. Experiments demonstrate that MoonCast outperforms baselines, with particularly notable improvements in spontaneity and coherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。