通过多尺度声学与韵律一致性提升文本语音编辑的自然度
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
- 引入分层局部声学平滑约束,确保编辑区与未编辑区声学特征无缝衔接
- 设计对比全局韵律一致性约束,保持编辑后语音整体语调与原声一致
- 在VCTK和LibriTTS上优于现有方法,适合语音编辑与内容重制场景
文本语音编辑(TSE)允许用户通过直接修改对应文本来编辑语音而不改变原始录音。现有方法通常在训练中最小化生成语音与参考语音在编辑区域内的差异以实现流畅性。然而,编辑区域的生成语音应在局部和全局层面保持与未编辑区域及原始语音的声学与韵律一致性。为此,我们提出基于先前的FluentEditor模型的新型流畅语音编辑方案FluentEditor2,通过建模多尺度声学与韵律一致性训练准则。具体而言,针对局部声学一致性,提出分层局部声学平滑约束,对语音帧、音素和词在编辑区与未编辑区交界处的声学属性进行对齐;针对全局韵律一致性,提出对比全局韵律一致性约束,使编辑区语音与原句韵律保持一致。在VCTK和LibriTTS数据集上的大量实验表明,FluentEditor2在主观与客观评价上均优于现有神经网络基线方法,包括Editspeech、Campnet、A$^3$T、FluentSpeech及我们的FluentEditor。消融实验进一步验证了各模块对系统整体有效性的贡献。语音演示见:https://github.com/Ai-S2-Lab/FluentEditor2。
原文摘要 · Abstract (English)
Text-based speech editing (TSE) allows users to edit speech by modifying the corresponding text directly without altering the original recording. Current TSE techniques often focus on minimizing discrepancies between generated speech and reference within edited regions during training to achieve fluent TSE performance. However, the generated speech in the edited region should maintain acoustic and prosodic consistency with the unedited region and the original speech at both the local and global levels. To maintain speech fluency, we propose a new fluency speech editing scheme based on our previous \textit{FluentEditor} model, termed \textit{\textbf{FluentEditor2}}, by modeling the multi-scale acoustic and prosody consistency training criterion in TSE training. Specifically, for local acoustic consistency, we propose \textit{hierarchical local acoustic smoothness constraint} to align the acoustic properties of speech frames, phonemes, and words at the boundary between the generated speech in the edited region and the speech in the unedited region. For global prosody consistency, we propose \textit{contrastive global prosody consistency constraint} to keep the speech in the edited region consistent with the prosody of the original utterance. Extensive experiments on the VCTK and LibriTTS datasets show that \textit{FluentEditor2} surpasses existing neural networks-based TSE methods, including Editspeech, Campnet, A$^3$T, FluentSpeech, and our Fluenteditor, in both subjective and objective. Ablation studies further highlight the contributions of each module to the overall effectiveness of the system. Speech demos are available at: \url{https://github.com/Ai-S2-Lab/FluentEditor2}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。