arXiv:2510.23541eess.AScs.SD2025-10被引 11

SoulX-Podcast可生成90分钟以上多角色方言对话,自然度达顶尖水平。

SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity

  • 融合语气控制与多方言支持,实现多角色对话语音生成
  • 连续生成超90分钟对话,声线稳定、换人流畅
  • 适合个性化播客、虚拟主播等需要真实语调的场景

近年来文本转语音(TTS)技术显著提升了语音表现力与自然度,但多数系统仍局限于单说话人合成,难以生成连贯的多说话人对话。本文提出SoulX-Podcast系统,专为播客风格的多轮、多说话人对话语音生成设计,同时在传统单人语音合成任务中达到当前最优性能。为满足多轮口语对话更高的自然度要求,SoulX-Podcast集成多种副语言控制机制,支持普通话、英语及四川话、河南话、粤语等多种中文方言,实现更个性化的播客式语音生成。实验表明,该系统可连续生成超过90分钟的对话,保持稳定的说话人音色与平滑的说话人切换。此外,说话人表现出情境自适应的韵律特征,随对话推进呈现自然的节奏与语调变化。在多项评估指标下,SoulX-Podcast在单人语音合成与多轮对话语音合成任务中均达到领先水平。

原文摘要 · Abstract (English)

Recent advances in text-to-speech (TTS) synthesis have significantly improved speech expressiveness and naturalness. However, most existing systems are tailored for single-speaker synthesis and fall short in generating coherent multi-speaker conversational speech. This technical report presents SoulX-Podcast, a system designed for podcast-style multi-turn, multi-speaker dialogic speech generation, while also achieving state-of-the-art performance in conventional TTS tasks. To meet the higher naturalness demands of multi-turn spoken dialogue, SoulX-Podcast integrates a range of paralinguistic controls and supports both Mandarin and English, as well as several Chinese dialects, including Sichuanese, Henanese, and Cantonese, enabling more personalized podcast-style speech generation. Experimental results demonstrate that SoulX-Podcast can continuously produce over 90 minutes of conversation with stable speaker timbre and smooth speaker transitions. Moreover, speakers exhibit contextually adaptive prosody, reflecting natural rhythm and intonation changes as dialogues progress. Across multiple evaluation metrics, SoulX-Podcast achieves state-of-the-art performance in both monologue TTS and multi-turn conversational speech synthesis.

语音合成多说话人方言支持播客生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。