arXiv:2607.00363cs.SDcs.AI2026-07中稿 · INTERSPEECH 2026

用统一引导框架提升语音生成速度与音色保真度

Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis

论文配图:Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis
图 1 · 摘自论文原文
  • 通过异质数据增强和新型模型引导机制分离语言与声学特征
  • 推理速度提升近3倍,说话人相似度优于当前最优基线
  • 适合追求高效高保真语音合成的开发者与研究者

流匹配(Flow Matching, FM)已成为强大的语音生成范式,但受限于高推理延迟和音色泄露问题。为此,我们提出统一引导框架,通过两种互补策略提升生成效率与鲁棒性。在数据层面,引入异质数据增强(Heterogeneous Augmentation),促使模型将语言内容与声学残差解耦;在模型层面,提出增强型模型引导机制,融合轨迹修正与新型内在引导目标,将条件知识提炼至网络权重中,并规整推理轨迹路径,从而消除无分类器引导(Classifier-Free Guidance, CFG)开销。实验表明,该框架使推理速度提升近三倍,同时在说话人相似度上显著优于当前最优基线。

原文摘要 · Abstract (English)

Flow Matching (FM) has emerged as a powerful paradigm for speech generation but remains constrained by high inference latency and timbre leakage. To address these bottlenecks, we propose a unified guidance framework that enhances generation efficiency and robustness through two complementary strategies. On the data front, we introduce Data-guidance via heterogeneous augmentation, encouraging the model to disentangle linguistic content from acoustic residue. In parallel, we propose an enhanced Model-guidance mechanism that synergizes trajectory rectification with a novel intrinsic guidance objective. This approach distills conditional knowledge into network weights and straightens inference trajectory path, thereby eliminating Classifier-Free Guidance (CFG) overhead. Experiments demonstrate that our framework accelerates inference by nearly three times while effectively improving speaker similarity compared to state-of-the-art baselines.

语音合成流匹配高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。