arXiv:2507.17208eess.AS2025-07中稿 · INTERSPEECH 2025

融合自监督学习与数字信号处理,提升语音基频估计精度

SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch

  • 利用数字信号处理生成先验基频分布,结合梯度优化绝对基频
  • 在标准数据集上优于传统方法,基频估计误差显著降低
  • 适合语音分析、声学建模等需要高精度基频的场景

我们提出SLASH,一种基于自监督学习(SSL)的语音基频估计方法。为克服传统SSL方法依赖相对基频差异的局限,该方法通过引入源自数字信号处理(DSP)的先验基频分布,并利用目标谱图与可微分DSP生成谱图之间的损失进行梯度下降优化,实现绝对基频的精准估计。为稳定优化过程,采用新型谱图生成方法,跳过复杂的波形重建步骤。此外,通过可微分DSP准确预测语音中的非周期成分,提升了方法在语音信号处理中的适用性。实验表明,所提方法在性能上超越基准的DSP与基于SSL的基频估计方法,归因于SSL与DSP的有效融合。

原文摘要 · Abstract (English)

We present SLASH, a pitch estimation method of speech signals based on self-supervised learning (SSL). To enhance the performance of conventional SSL-based approaches that primarily depend on the relative pitch difference derived from pitch shifting, our method incorporates absolute pitch values by 1) introducing a prior pitch distribution derived from digital signal processing (DSP), and 2) optimizing absolute pitch through gradient descent with a loss between the target and differentiable DSP-derived spectrograms. To stabilize the optimization, a novel spectrogram generation method is used that skips complicated waveform generation. In addition, the aperiodic components in speech are accurately predicted through differentiable DSP, enhancing the method's applicability to speech signal processing. Experimental results showed that the proposed method outperformed both baseline DSP and SSL-based pitch estimation methods, attributed to the effective integration of SSL and DSP.

语音处理自监督学习基频估计数字信号处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。