提出新调度方法与修正机制,提升零样本语音合成质量。
Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech
- 基于费雪-劳伊速度设计最优调度路径,无需训练
- 引入有限步数概率修正,降低生成误差
- 在语音自然度和说话人相似性上均领先于现有模型
度量诱导的离散流匹配(MI-DFM)利用词元-隐变量几何结构进行离散生成,但其实际应用受限于两个问题:启发式调度器需超参数搜索,以及一阶连续时间马尔可夫链(CTMC)求解器带来的有限步数路径追踪误差。本文提出两种改进:首先,推导出适用于指定标量参数化概率路径的动能最优调度器,并在MI-DFM中实现为无需训练的数值调度,使路径以恒定费雪-劳伊速度遍历;其次,引入有限步数矩修正机制,在保持CTMC跳转目标分布的前提下调整跳跃概率。所提出的GibbsTTS方法在基于编解码器的零样本文本到语音任务中进行了验证。在统一架构与大规模数据集上的对比实验表明,GibbsTTS在客观自然度上表现最佳,且在主观评价中优于掩码式离散生成基线模型。此外,与评估的先进语音合成系统相比,GibbsTTS在四个测试集中的三个达到最高说话人相似性,第四组排名第二。
原文摘要 · Abstract (English)
Metric-induced discrete flow matching (MI-DFM) exploits token-latent geometry for discrete generation, but its practical use is limited by two issues: heuristic schedulers requiring hyperparameter search, and finite-step path-tracking error from its first-order continuous-time Markov chain (CTMC) solver. We address both issues. First, we derive a kinetic-optimal scheduler for prescribed scalar-parameterized probability paths, and instantiate it for MI-DFM as a training-free numerical schedule that traverses the path at constant Fisher-Rao speed. Second, we introduce a finite-step moment correction that adjusts the jump probability while preserving the CTMC jump destination distribution. We validate the resulting method, GibbsTTS, on codec-based zero-shot text-to-speech (TTS). Under controlled comparisons with a unified architecture and large-scale dataset, GibbsTTS achieves the best objective naturalness and is preferred in subjective evaluations over masked discrete generative baselines. Additionally, in comparison with the evaluated state-of-the-art TTS systems, GibbsTTS shows strong speaker similarity, achieving the highest similarity on three of four test sets and ranking second on the fourth. Project page: https://ydqmkkx.github.io/GibbsTTSProject
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。