arXiv:2509.17516eess.AS2025-09被引 1

让语音合成更连贯可控,专为多角色有声书设计

Audiobook-CC: Controllable Long-context Speech Generation for Multicast Audiobook

  • 引入上下文机制确保长文本一致性
  • 分离风格与语义控制,提升表达精准度
  • 自蒸馏增强情感表现力,适合有声书制作

现有文本转语音系统主要聚焦单句合成,缺乏充分的上下文建模和细粒度性能控制能力,难以生成连贯的多角色有声书。为此,我们提出一种面向多播有声书的上下文感知、情绪可控语音合成框架,包含三项核心创新:用于保持上下文一致性的上下文机制,将风格控制与语音提示解耦以保障语义一致性的解纠缠范式,以及通过自蒸馏提升情感表现力与指令可控性的方法。实验表明,该框架在旁白、对话及整章生成任务中均显著优于现有基线模型。消融实验证实了各模块的有效性。演示样本可在 https://everest-ai.github.io/ 查看。

原文摘要 · Abstract (English)

Existing text-to-speech systems predominantly focus on single-sentence synthesis and lack adequate contextual modeling as well as fine-grained performance control capabilities for generating coherent multicast audiobooks. To address these limitations, we propose a context-aware and emotion controllable speech synthesis framework specifically engineered for multicast audiobooks with three key innovations: a context mechanism for contextual consistency, a disentanglement paradigm to decouple style control from speech prompts for semantic consistency, and self-distillation to boost emotional expressiveness and instruction controllability. Experimental results show superior performance across the generation of narration, dialogue, and the whole chapter, significantly outperforming existing baselines. Ablation studies are conducted to validate the effectiveness of our proposed methods. Demo samples can be found in https://everest-ai.github.io/.

语音合成有声书长文本生成情绪控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。