一个支持多说话人语音与音频生成的统一模型,能零样本和指令控制地创作声音。
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

- 通过统一MoE架构实现多任务多模态建模,支持零样本与指令驱动生成。
- 在零样本和指令任务上均超越现有方法,表达力得分最优。
- 适合动画配音、有声剧等需自定义声音的场景创作者使用。
语音与音频生成广泛应用于动画配音、有声剧、电影、广告、游戏、播客及短视频制作。创作者常需无参考录音设计声音,用自然语言控制说话风格,结合环境音效,并后续复用已设计的声音。为此,需同时支持零样本和指令任务的多说话人语音与音频生成。指令任务要求提供环境描述、说话风格及细粒度内容提示,而零样本任务则依赖参考音频与相同内容。本文从数据与模型双方面入手:提出SwanData-Caption,清洗原始数据,增加合成覆盖,标注多层次准确描述;并构建SwanTale模型,采用SwanVAE实现高质量多模态生成,引入奖励条件质量控制与Engram条件化,并结合统一MoE处理多任务与多音频模态。此外,通过课程学习与GRPO后训练,使模型逐步提升能力。实验表明,SwanTale在多个关键零样本与指令指标上领先,两类任务中表达力得分最佳,支持包含多说话人的复杂指令生成。演示见 https://swanaigc.github.io/#swantale。
原文摘要 · Abstract (English)
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, support acoustic scenes with environments and audio effects, and later reuse the designed voices. Therefore, it is important to support multi-speaker speech and audio generation for both instruct and zero-shot tasks. The instruct task requires a caption of the environment, speaker styles, and fine-grained content, while the zero-shot task uses reference audio together with the same fine-grained content. We address these tasks from both the data and model sides. First, we propose SwanData-Caption, which cleans raw speech and audio data, adds targeted synthetic coverage, and annotates diverse and accurate multi-level captions. Then, we propose SwanTale, a multi-speaker expressive speech and audio generation model that supports both zero-shot and instruct tasks. We introduce SwanVAE to support high-quality multi-audio-modality generation. Then, we adopt reward-conditioned quality control and Engram conditioning, along with Unified MoE for multi-task and multi-audio-modality modeling. In addition, we use curriculum learning and GRPO post-training to let the model progressively learn and strengthen its capabilities. Experimental results show that SwanTale leads on multiple key zero-shot and instruct metrics, achieves the best expressiveness scores in both tasks, and supports complex instruct generation involving multi-speaker speech and audio. Demos can be found at https://swanaigc.github.io/#swantale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。