用Wave-U-Net增强基音波形,实现高精度瞬时音高估计
Instantaneous Pitch Estimation via Wave-U-Net-Based Fundamental Waveform Enhancement
- 将基音提取视为语音增强问题,用Wave-U-Net重建基音波形
- 在语音、歌唱、乐器等多场景下均优于传统方法
- 特别适合噪声大或音高变化剧烈的信号分析
瞬时音高估计在分析快速音高变化(如语调和演唱技巧)中至关重要。传统方法需先从含谐波和噪声的信号中分离基音波形,其精度易受基音滤波不完善的干扰。本文将基音波形过滤建模为语音增强任务,训练Wave-U-Net模型从输入语音中提取基音波形,再通过估计基音波形解析信号的瞬时频率获取瞬时音高。实验表明,该方法在语音、歌唱、乐器及降质语音等多种场景下均显著优于传统确定性方法,实现了准确且鲁棒的瞬时音高估计。
原文摘要 · Abstract (English)
Instantaneous pitch estimation plays an important role in analyzing steep pitch variations such as speech prosody and singing techniques. Conventional approaches estimate instantaneous frequency after isolating the fundamental waveform from signals that contain harmonics and noise, which makes the accuracy sensitive to imperfect fundamental filtering. In this study, we formulate fundamental waveform filtering as a speech enhancement problem. Specifically, we train a Wave-U-Net model to extract a fundamental waveform from an input speech signal. The instantaneous pitch is then obtained by computing the instantaneous frequency from the analytic signal of the estimated fundamental waveform. Experimental results show that the proposed method outperforms conventional deterministic approaches and provides accurate and robust instantaneous pitch estimation across diverse domains, including speech, singing voice, musical instruments, and degraded speech signals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。