OZSpeech实现一键零样本语音合成,提升克隆精度与自然度。
OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching
- 采用学习先验条件流匹配,一步完成采样,无需迭代
- 在内容准确率、语调生成和声线保留上优于现有方法
- 适合需要快速高质量语音克隆的场景
近年来,文本转语音(TTS)系统因深度学习和神经网络架构的进步而取得显著进展。以往方法通常将语音视为数据分布,在流匹配框架中使用波形或频谱图等传统表示,但存在忽略语音属性多样性和训练引入额外约束导致计算成本高等问题。为此,我们提出OZSpeech,首个探索最优传输条件流匹配的单步零样本语音合成方法,采用学习先验作为条件,跳过前序状态,减少采样步骤。该方法在离散化语音特征的分量上操作,实现各语音属性的解耦建模,从而增强对提示语音的精准克隆能力。实验表明,该方法在内容准确性、自然度、韵律生成及说话人风格保留方面均优于现有方法。音频样例见演示页:https://ozspeech.github.io/OZSpeech_Web/
原文摘要 · Abstract (English)
Text-to-speech (TTS) systems have seen significant advancements in recent years, driven by improvements in deep learning and neural network architectures. Viewing the output speech as a data distribution, previous approaches often employ traditional speech representations, such as waveforms or spectrograms, within the Flow Matching framework. However, these methods have limitations, including overlooking various speech attributes and incurring high computational costs due to additional constraints introduced during training. To address these challenges, we introduce OZSpeech, the first TTS method to explore optimal transport conditional flow matching with one-step sampling and a learned prior as the condition, effectively disregarding preceding states and reducing the number of sampling steps. Our approach operates on disentangled, factorized components of speech in token format, enabling accurate modeling of each speech attribute, which enhances the TTS system's ability to precisely clone the prompt speech. Experimental results show that our method achieves promising performance over existing methods in content accuracy, naturalness, prosody generation, and speaker style preservation. Audio samples are available at our demo page https://ozspeech.github.io/OZSpeech_Web/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。