arXiv:2505.19384cs.CLcs.SD2025-05

无需训练新说话人,用渐进式风格适配实现高质量语音合成

GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor

  • 通过分层编码逐步提取声学参考中的说话风格
  • 在未见说话人上实现自然度、相似度和可懂度的优秀表现
  • 结构可解释且可控,适合需要灵活调节语音风格的场景

我们提出渐进式风格适配语音合成(GSA-TTS),其新颖的风格编码器能从声学参考中逐步编码说话风格,实现零样本语音合成。GSA首先捕获每个语义音素的局部风格,再通过自注意力机制整合为全局风格条件。这种语义与分层编码策略为声学模型提供了鲁棒且丰富的风格表征。我们在未见过的说话人上测试了GSA-TTS,结果在自然度、说话人相似度和可懂度方面表现优异。此外,其分层结构还带来了可解释性与可控性的潜力。

原文摘要 · Abstract (English)

We present the gradual style adaptor TTS (GSA-TTS) with a novel style encoder that gradually encodes speaking styles from an acoustic reference for zero-shot speech synthesis. GSA first captures the local style of each semantic sound unit. Then the local styles are combined by self-attention to obtain a global style condition. This semantic and hierarchical encoding strategy provides a robust and rich style representation for an acoustic model. We test GSA-TTS on unseen speakers and obtain promising results regarding naturalness, speaker similarity, and intelligibility. Additionally, we explore the potential of GSA in terms of interpretability and controllability, which stems from its hierarchical structure.

语音合成零样本风格迁移可控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。