arXiv:2509.15140cs.SDcs.CL2025-09被引 4

提出快速鲁棒的语音基频估计模型,噪声下性能优越且推理极快。

FCPE: A Fast Context-based Pitch Estimation Model

  • 基于深度可分离卷积的Lynx-Net架构,高效捕捉梅尔频谱特征。
  • 在MIR-1K数据集上达96.79%原始基频准确率,与顶尖方法相当。
  • 单卡RTX 4090实测实时因子仅0.0062,适合低延迟应用场景。

单音音频中的基频估计对MIDI转录和歌声转换至关重要,但现有方法在噪声环境下性能显著下降。本文提出FCPE,一种基于上下文的快速基频估计模型,采用Lynx-Net架构结合深度可分离卷积,在有效捕捉梅尔频谱特征的同时保持低计算开销和强噪声容忍能力。实验表明,该方法在MIR-1K数据集上达到96.79%的原始基频准确率(RPA),与当前最优方法持平。在单块RTX 4090 GPU上实时因子(RTF)仅为0.0062,显著优于现有算法的效率。代码已开源:https://github.com/CNChTu/FCPE。

原文摘要 · Abstract (English)

Pitch estimation (PE) in monophonic audio is crucial for MIDI transcription and singing voice conversion (SVC), but existing methods suffer significant performance degradation under noise. In this paper, we propose FCPE, a fast context-based pitch estimation model that employs a Lynx-Net architecture with depth-wise separable convolutions to effectively capture mel spectrogram features while maintaining low computational cost and robust noise tolerance. Experiments show that our method achieves 96.79\% Raw Pitch Accuracy (RPA) on the MIR-1K dataset, on par with the state-of-the-art methods. The Real-Time Factor (RTF) is 0.0062 on a single RTX 4090 GPU, which significantly outperforms existing algorithms in efficiency. Code is available at https://github.com/CNChTu/FCPE.

基频估计语音处理轻量模型实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。