arXiv:2606.25672eess.AScs.SD2026-06

提出新方法提升零样本语音合成的说话人相似度与文本准确率平衡。

Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS

论文配图:Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS
图 1 · 摘自论文原文
  • 分解引导场为文本、说话人和联合残差,独立控制
  • 在F5-TTS和CosyVoice2上提升说话人相似度,文本正确率仍领先
  • 适合需要高保真说话人还原的零样本语音合成场景

Classifier-free guidance(CFG)广泛用于基于流匹配的零样本文语合成(TTS),生成通常由目标文本和提示语音信号两种条件控制。标准CFG联合强化这些条件,而近期分支选择性引导方法尝试分别增强文本或说话人条件,常导致文本准确性与说话人相似度之间的权衡。本文重新审视在独立掩码文本与语音提示条件下的CFG,将引导场分解为文本、说话人和联合残差。我们发现传统说话人选择性引导会将说话人残差与联合残差纠缠,可能干扰文本相关生成。基于此观察,提出联合残差重加权,在标准CFG框架内独立控制说话人与联合残差。在F5-TTS和CosyVoice2上的实验表明,该方法在保持竞争性文本正确率的同时提升了说话人相似度,证明联合残差在平衡说话人保真度与文本准确性方面的有效性。

原文摘要 · Abstract (English)

Classifier-free guidance (CFG) is widely used in flow-matching-based zero-shot text-to-speech (TTS), where generation is typically controlled by two conditions: the target text and a prompt speech signal. Standard CFG strengthens these conditions jointly, while recent branch-selective guidance methods attempt to enhance text or speaker conditioning separately, often leading to a trade-off between text correctness and speaker similarity. In this paper, we revisit the CFG under independently masked text and speech-prompt conditions, and decompose the guidance field into text, speaker, and joint residuals. We show that conventional speaker-selective guidance entangles the speaker residual with the joint residual, which may disturb text-related generation. Based on this observation, we propose joint residual reweighting, which independently controls the speaker and joint residuals within the standard CFG framework. Experiments on F5-TTS and CosyVoice2 show that the proposed method improves speaker similarity while maintaining competitive text correctness, demonstrating the usefulness of the joint residual for balancing speaker fidelity and text accuracy in zero-shot TTS.

零样本语音合成流匹配说话人控制引导生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。