arXiv:2509.19668eess.AScs.AI2025-09被引 1

提出分阶段引导策略,提升零样本语音合成的音色保真度与文本一致性。

Selective Classifier-free Guidance for Zero-shot Text-to-speech

  • 早期用标准引导,后期切换为选择性引导,平衡音色与文本表现
  • 在英汉双语实验中发现引导效果受文本表示方式显著影响
  • 首次揭示语音合成中条件分离引导的语义依赖性,适合语音生成研究者

在零样本文本到语音合成中,如何在保持目标说话人特征与忠实于文本内容之间取得平衡仍是挑战。尽管分类器无关引导(CFG)在图像生成中表现优异,其在语音合成中的应用仍不充分。通过分离用于引导的条件,可在语音合成中实现不同特性的权衡。本文评估了原本为图像生成设计的CFG策略在语音合成中的适应性,并拓展了分离条件下的CFG方法。结果表明,图像生成中有效的CFG策略在语音合成中通常无效。我们发现,若在早期阶段使用标准CFG,后期仅在关键阶段采用选择性CFG,可提升说话人相似度且限制文本一致性的下降。令人意外的是,选择性CFG的效果高度依赖文本表示方式:即使使用相同模型,英语与汉语的表现差异显著。

原文摘要 · Abstract (English)

In zero-shot text-to-speech, achieving a balance between fidelity to the target speaker and adherence to text content remains a challenge. While classifier-free guidance (CFG) strategies have shown promising results in image generation, their application to speech synthesis are underexplored. Separating the conditions used for CFG enables trade-offs between different desired characteristics in speech synthesis. In this paper, we evaluate the adaptability of CFG strategies originally developed for image generation to speech synthesis and extend separated-condition CFG approaches for this domain. Our results show that CFG strategies effective in image generation generally fail to improve speech synthesis. We also find that we can improve speaker similarity while limiting degradation of text adherence by applying standard CFG during early timesteps and switching to selective CFG only in later timesteps. Surprisingly, we observe that the effectiveness of a selective CFG strategy is highly text-representation dependent, as differences between the two languages of English and Mandarin can lead to different results even with the same model.

语音合成引导生成零样本文本表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。