arXiv:2507.11233cs.SDeess.AS2025-07中稿 · ISMIR 2025被引 1

用专用声学特征提升音高估计,更准更省参数

Improving Neural Pitch Estimation with SWIPE Kernels

  • 用锯齿波启发的SWIPE特征做音频前端
  • 网络规模缩小十倍,性能不降反升
  • 适合追求高效音高估计的开发者

神经网络已成为高精度音高与周期性估计的主流方法。尽管研究广泛聚焦于网络架构和训练方式改进,但多数方法直接处理原始音频波形或通用时频表示。本文研究了锯齿波启发的音高估计(SWIPE)核作为音频前端的效果,发现这种手工设计的任务特异性特征可显著提升神经音高估计算法的准确性、抗噪能力及参数效率。在常见数据集上评估了监督与自监督的前沿模型,结果表明使用SWIPE前端可使网络规模缩小一个数量级而性能不下降。此外,SWIPE算法本身也表现优异,其精度超过当前主流自监督神经音高估计器。

原文摘要 · Abstract (English)

Neural networks have become the dominant technique for accurate pitch and periodicity estimation. Although a lot of research has gone into improving network architectures and training paradigms, most approaches operate directly on the raw audio waveform or on general-purpose time-frequency representations. We investigate the use of Sawtooth-Inspired Pitch Estimation (SWIPE) kernels as an audio frontend and find that these hand-crafted, task-specific features can make neural pitch estimators more accurate, robust to noise, and more parameter-efficient. We evaluate supervised and self-supervised state-of-the-art architectures on common datasets and show that the SWIPE audio frontend allows for reducing the network size by an order of magnitude without performance degradation. Additionally, we show that the SWIPE algorithm on its own is much more accurate than commonly reported, outperforming state-of-the-art self-supervised neural pitch estimators.

音高估计前端设计参数效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。