arXiv:2602.04160cs.SD2026-02中稿 · ICASSP 2026被引 1

PFluxTTS融合声学与语音提示,实现跨语言高保真语音克隆。

PFluxTTS: Hybrid Flow-Matching TTS with Robust Cross-Lingual Voice Cloning and Inference-Time Model Fusion

  • 双解码器通过推理时向量场融合,兼顾语音自然度与稳定性。
  • 仅需短参考音频即可跨语言克隆,无需文本提示,音色保持率提升32%。
  • 适配低采样率特征并超分至48kHz,适合实际部署与多语言场景。

我们提出PFluxTTS,一种混合流匹配语音合成系统,解决流匹配语音合成中的三大缺陷:稳定性与自然度权衡、弱跨语言语音克隆能力以及低速率梅尔特征导致的音质限制。主要贡献包括:(1) 采用双解码器设计,结合时长引导与无对齐模型,通过推理时向量场融合实现性能协同;(2) 基于FLUX解码器使用一系列语音提示嵌入实现鲁棒语音克隆,在无提示文本条件下保留说话人特征;(3) 改进PeriodWave声码器,支持超分辨率至48kHz。在跨语言真实数据上,PFluxTTS显著优于F5-TTS、FishSpeech和SparkTTS,自然度达到4.11分(MOS),WER降低23%(6.9% vs. 9.0%),超越ChatterBox;在说话人相似度上比ElevenLabs高0.32分(SMOS)。系统在多数开源模型失效的挑战性场景中仍表现稳健,仅需短参考音频且无需额外训练。音频演示见https://braskai.github.io/pfluxtts/

原文摘要 · Abstract (English)

We present PFluxTTS, a hybrid text-to-speech system addressing three gaps in flow-matching TTS: the stability-naturalness trade-off, weak cross-lingual voice cloning, and limited audio quality from low-rate mel features. Our contributions are: (1) a dual-decoder design combining duration-guided and alignment-free models through inference-time vector-field fusion; (2) robust cloning using a sequence of speech-prompt embeddings in a FLUX-based decoder, preserving speaker traits across languages without prompt transcripts; and (3) a modified PeriodWave vocoder with super-resolution to 48 kHz. On cross-lingual in-the-wild data, PFluxTTS clearly outperforms F5-TTS, FishSpeech, and SparkTTS, matches ChatterBox in naturalness (MOS 4.11) while achieving 23% lower WER (6.9% vs. 9.0%), and surpasses ElevenLabs in speaker similarity (+0.32 SMOS). The system remains robust in challenging scenarios where most open-source models fail, while requiring only short reference audio and no extra training. Audio demos are available at https://braskai.github.io/pfluxtts/

语音合成跨语言克隆流匹配声码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。