arXiv:2608.11737cs.SD2026-08

用流匹配联合训练语音分词与合成,提升音质和说话人相似度。

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization

论文配图:Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization
图 1 · 摘自论文原文
  • 分词器与声学模型联合训练,直接优化生成任务
  • 110K小时数据训练,语音可懂度优于真人录音
  • 无需微调即可用于零样本语音转换,适合语音合成与转换场景

当前零样本文本到语音系统中,传统语义分词器通常通过有监督语音识别或自监督学习目标进行优化。但由于语音固有的特性,语义与声学信息难以完全解耦;基于ASR的分词器为专注语言内容而丢弃声学细节,导致模型难以实现最佳说话人相似度。此外,这些分词器独立训练,缺乏对下游声学生成任务的直接监督,造成提取的离散标记与声学模型所需的连续空间之间存在特征鸿沟,从根本上限制了合成质量上限。为此,我们提出Phoenix TTS,一个将表征学习与生成声学建模紧密耦合的统一框架。具体而言,我们的语音分词器在重建自监督特征以保持语义丰富性的同时,还接收来自流匹配训练损失的直接监督。通过这种联合训练范式,提取的离散标记成功保留了关键语义信息,并天然对齐下游流匹配模型的特征空间。全面评估表明,该系统高效且有效:在110,000小时数据上训练,语音可懂度优异,其字错误率(WER)持续低于真实录音。同时,保持强大的零样本说话人相似度,媲美或超越多个主流大规模基线。此外,作为这一联合训练的有益副产品,所学分词器可无缝应用于零样本语音转换任务,无需任务特定微调。

原文摘要 · Abstract (English)

In current zero-shot text-to-speech systems, conventional semantic tokenizers are typically optimized using supervised automatic speech recognition or self-supervised learning objectives. However, due to the inherent nature of speech, semantic and acoustic information cannot be completely decoupled, and ASR-based tokenizers discard acoustic details to focus on linguistic content; models relying on them usually struggle to achieve optimal speaker similarity. Furthermore, these tokenizers are optimized independently and lack direct supervision from downstream acoustic generation tasks. This isolated training creates a feature gap between the extracted discrete tokens and the continuous space required by acoustic models, fundamentally bottlenecking the upper bound of synthesis quality. To bridge this gap, we propose Phoenix TTS, a unified framework that tightly couples representation learning with generative acoustic modeling. Specifically, our speech tokenizer is optimized to reconstruct self-supervised features to maintain semantic richness, while simultaneously receiving direct supervision from a Flow Matching training loss. Through this joint training paradigm, the extracted discrete tokens successfully preserve essential semantic information and natively align with the feature space of the downstream Flow Matching model. Comprehensive evaluations highlight the efficiency and effectiveness of Phoenix TTS. Trained on 110K hours of data, the system achieves excellent speech intelligibility, yielding WER that consistently falls below that of ground-truth recordings. Simultaneously, it maintains robust zero-shot speaker similarity, rivaling or outperforming several prominent large-scale baselines. Furthermore, as an advantageous byproduct of this unified training, the learned tokenizer can be seamlessly adapted to zero-shot voice conversion tasks without requiring task-specific fine-tuning.

语音合成流匹配零样本转换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。