用统一引导框架提升语音生成速度与音色保真度
Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis

- 通过异质数据增强和新型模型引导机制分离语言与声学特征
- 推理速度提升近3倍,说话人相似度优于当前最优基线
- 适合追求高效高保真语音合成的开发者与研究者
流匹配(Flow Matching, FM)已成为强大的语音生成范式,但受限于高推理延迟和音色泄露问题。为此,我们提出统一引导框架,通过两种互补策略提升生成效率与鲁棒性。在数据层面,引入异质数据增强(Heterogeneous Augmentation),促使模型将语言内容与声学残差解耦;在模型层面,提出增强型模型引导机制,融合轨迹修正与新型内在引导目标,将条件知识提炼至网络权重中,并规整推理轨迹路径,从而消除无分类器引导(Classifier-Free Guidance, CFG)开销。实验表明,该框架使推理速度提升近三倍,同时在说话人相似度上显著优于当前最优基线。
原文摘要 · Abstract (English)
Flow Matching (FM) has emerged as a powerful paradigm for speech generation but remains constrained by high inference latency and timbre leakage. To address these bottlenecks, we propose a unified guidance framework that enhances generation efficiency and robustness through two complementary strategies. On the data front, we introduce Data-guidance via heterogeneous augmentation, encouraging the model to disentangle linguistic content from acoustic residue. In parallel, we propose an enhanced Model-guidance mechanism that synergizes trajectory rectification with a novel intrinsic guidance objective. This approach distills conditional knowledge into network weights and straightens inference trajectory path, thereby eliminating Classifier-Free Guidance (CFG) overhead. Experiments demonstrate that our framework accelerates inference by nearly three times while effectively improving speaker similarity compared to state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。