一歩で非並列音声変換を実現、高速かつ高品質な変換が可能。
MeanVoiceFlow: One-step Nonparallel Voice Conversion with Mean Flows
- 平均速度を使った単一ステップ推論で、逐次処理の遅さを克服。
- 学習時にゼロ入力制約を導入し、モデルの安定性を向上。
- 事前学習不要で初期から訓練可能、実用性が高い。
在语音转换(VC)应用中,扩散模型和流匹配模型表现出优异的语音质量和说话人相似度性能。然而,由于其迭代推理过程缓慢而受限。为此,我们提出一种基于平均流的新颖单步非并行语音转换模型——MeanVoiceFlow,可从零开始训练,无需预训练或蒸馏。与传统流匹配使用瞬时速度不同,平均流采用平均速度更准确地计算单步推理路径上的时间积分。但训练平均速度需要其导数来计算目标速度,可能导致不稳定性。因此,我们引入结构边缘重建损失作为零输入约束,适度正则化模型的输入-输出行为,避免有害的统计平均。此外,提出条件扩散输入训练,训练和推理时均使用噪声与源数据的混合输入,使模型有效利用源信息的同时保持训练与推理的一致性。实验结果验证了这些技术的有效性,表明即使从零开始训练,MeanVoiceFlow 的性能仍可媲美先前多步及蒸馏基模型。音频样例见 https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/meanvoiceflow/。
原文摘要 · Abstract (English)
In voice conversion (VC) applications, diffusion and flow-matching models have exhibited exceptional speech quality and speaker similarity performances. However, they are limited by slow conversion owing to their iterative inference. Consequently, we propose MeanVoiceFlow, a novel one-step nonparallel VC model based on mean flows, which can be trained from scratch without requiring pretraining or distillation. Unlike conventional flow matching that uses instantaneous velocity, mean flows employ average velocity to more accurately compute the time integral along the inference path in a single step. However, training the average velocity requires its derivative to compute the target velocity, which can cause instability. Therefore, we introduce a structural margin reconstruction loss as a zero-input constraint, which moderately regularizes the input-output behavior of the model without harmful statistical averaging. Furthermore, we propose conditional diffused-input training in which a mixture of noise and source data is used as input to the model during both training and inference. This enables the model to effectively leverage source information while maintaining consistency between training and inference. Experimental results validate the effectiveness of these techniques and demonstrate that MeanVoiceFlow achieves performance comparable to that of previous multi-step and distillation-based models, even when trained from scratch. Audio samples are available at https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/meanvoiceflow/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。