arXiv:2511.09995eess.AS2025-11

提升零样本语音合成的说话人相似度,让声音更像目标说话人。

Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTS

  • 根据时间步与网络层动态调整说话人特征对齐方式。
  • 在研究与工业级数据集上显著提升说话人相似度。
  • 适用于多种语音合成架构,泛化能力强。

基于流匹配(Flow-Matching, FM)的零样本文语转换(TTS)系统具备高质量语音合成与强泛化能力。然而,此类系统在说话人表征方面的能力仍待深入探索,主要因FM框架缺乏显式的说话人监督信号。为此,我们对说话人信息分布进行了实证分析,发现其在时间步与网络层间呈现非均匀分布,凸显了自适应说话人对齐的必要性。据此,我们提出时间-层级自适应说话人对齐(Time-Layer Adaptive Speaker Alignment, TLA-SA),通过联合利用时间与层级变化增强说话人一致性。实验表明,TLA-SA在研究与工业级数据集上均显著提升基线系统的说话人相似度,并在包括仅解码器语言模型(LM)与自由式TTS等多种架构中具有良好泛化能力。

原文摘要 · Abstract (English)

Flow-Matching (FM)-based zero-shot text-to-speech (TTS) systems exhibit high-quality speech synthesis and robust generalization capabilities. However, the speaker representation ability of such systems remains underexplored, primarily due to the lack of explicit speaker-specific supervision in the FM framework. To this end, we conduct an empirical analysis of speaker information distribution and reveal its non-uniform allocation across time steps and network layers, underscoring the need for adaptive speaker alignment. Accordingly, we propose Time-Layer Adaptive Speaker Alignment (TLA-SA), a strategy that enhances speaker consistency by jointly leveraging temporal and hierarchical variations. Experimental results show that TLA-SA substantially improves speaker similarity over baseline systems on both research- and industrial-scale datasets and generalizes well across diverse model architectures, including decoder-only language model (LM)-based and free TTS systems. A demo is provided.

语音合成零样本TTS说话人对齐流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。