重新审视Vocos的相位重建问题,发现其音质差距源于相位建模不足。
Revisiting Vocos: That Phasiness Business in Time-Frequency Neural Vocoding
- 从相位重建角度重访Vocos,分析其在频域建模中的瓶颈。
- 量化结果显示频域声码器与时域声码器在带限梅尔谱输入下存在明显质量差距。
- 改进骨干网络以预测相位差,揭示1D卷积层阻碍了相位精确重建。
近期,时频域神经声码器已接近时域神经声码器的顶尖性能。Vocos是其中的典型代表,因其高效性备受关注,但其音频质量仍落后于时域声码器,原因尚存争议。为此,本文从相位重建视角重新审视Vocos。首先,使用带限梅尔谱作为输入,量化了时域与频域声码器之间的性能差距。随后,通过消融实验验证:Vocos架构对幅度建模有效,但对相位建模能力较弱。进一步地,我们将Vocos主干网络改造为预测相位差(相位重建的前驱),发现1D卷积层限制了其准确预测能力。研究指出,未来工作需关注能更好建模语音时频结构的归纳偏置,同时不牺牲对任意输入表示的支持。
原文摘要 · Abstract (English)
Recently, time-frequency neural vocoders have been approaching the state-of-the-art quality of time-domain neural vocoders. Vocos is a notable example due to its efficiency, but its audio quality lags behind the time-domain vocoders and the reasons remain debated. Thus, in this study, we revisit Vocos from a phase reconstruction perspective. First, we quantify the gap between time-domain and time-frequency domain vocoders using bandlimited mel spectrograms as inputs. Later, via an ablation study, we verify the Vocos architecture is effective for magnitude modeling, but less so for phase. We then adapt the Vocos backbone to predict phase differences, a precursor for phase reconstruction, and identify 1D convolutional layers are hindering their accurate prediction. Our findings indicate that future research needs to focus on inductive biases that allow the architecture to better model the time-frequency structure of speech signals, without sacrificing the support for arbitrary input representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。