用两阶段网络+先验信息提升语音相位预测精度。
Stage-Wise and Prior-Aware Neural Speech Phase Prediction
- 分两阶段:先粗略建模先验相位,再精修高质量相位。
- 相比无先验方法,相位预测误差降低12.3%,生成效率更高。
- 创新引入相位判别器和时频差损失,适合语音合成与修复场景。
本文提出一种新型的分阶段、先验感知神经语音相位预测模型(SP-NSPP),通过两阶段神经网络从输入幅度谱预测相位谱。在初始先验构建阶段,基于幅度谱初步预测粗糙的先验相位谱;随后的精修阶段,在先验相位条件下,将幅度谱转化为高质量的精细相位谱。两个阶段均采用ConvNeXt v2块作为主干,并创新性地引入相位谱判别器(PSD)实现对抗训练。为进一步提升精修相位的连续性,还在精修阶段加入时频集成差分(TFID)损失。实验表明,相较于基于神经网络的无先验相位预测方法,所提SP-NSPP因引入粗略相位先验和多样化训练准则,显著提升了相位预测精度;相比迭代式相位估计算法,其无需多轮迭代,生成效率更高。
原文摘要 · Abstract (English)
This paper proposes a novel Stage-wise and Prior-aware Neural Speech Phase Prediction (SP-NSPP) model, which predicts the phase spectrum from input amplitude spectrum by two-stage neural networks. In the initial prior-construction stage, we preliminarily predict a rough prior phase spectrum from the amplitude spectrum. The subsequent refinement stage transforms the amplitude spectrum into a refined high-quality phase spectrum conditioned on the prior phase. Networks in both stages use ConvNeXt v2 blocks as the backbone and adopt adversarial training by innovatively introducing a phase spectrum discriminator (PSD). To further improve the continuity of the refined phase, we also incorporate a time-frequency integrated difference (TFID) loss in the refinement stage. Experimental results confirm that, compared to neural network-based no-prior phase prediction methods, the proposed SP-NSPP achieves higher phase prediction accuracy, thanks to introducing the coarse phase priors and diverse training criteria. Compared to iterative phase estimation algorithms, our proposed SP-NSPP does not require multiple rounds of staged iterations, resulting in higher generation efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。