arXiv:2602.01727cs.SD2026-02中稿 · ICASSP 2026被引 1

改进投票法音高估计,提升准确性和抗噪能力

Voting-based Pitch Estimation with Temporal and Frequential Alignment and Correlation Aware Selection

  • 通过时频对齐和相关性筛选,优化多个音高估计器的投票机制
  • 在多种语音音乐数据上,干净环境下超越单个顶尖方法
  • 适合需要高鲁棒性的音高估计场景,如语音识别、音乐信息检索

投票法是一种用于基频估计的集成方法,虽实证显示其稳健性,但缺乏深入研究。本文提供该方法的有效性理论基础,解释了基频估计误差方差的降低,并引用康多塞陪审团定理说明其在有声/无声检测中的准确性。针对实际局限性,提出两项改进:1)预投票对齐流程,校正估计器间的时序与频率偏差;2)基于误差相关性的贪心算法,选取紧凑而高效的估计器子集。在涵盖语音、歌唱和音乐的多样化数据集上实验表明,采用对齐方法的方案在清洁条件下优于单个最先进的估计器,在噪声环境中仍保持稳健的有声/无声检测能力。

原文摘要 · Abstract (English)

The voting method, an ensemble approach for fundamental frequency estimation, is empirically known for its robustness but lacks thorough investigation. This paper provides a principled analysis and improvement of this technique. First, we offer a theoretical basis for its effectiveness, explaining the error variance reduction for fundamental frequency estimation and invoking Condorcet's jury theorem for voiced/unvoiced detection accuracy. To address its practical limitations, we propose two key improvements: 1) a pre-voting alignment procedure to correct temporal and frequential biases among estimators, and 2) a greedy algorithm to select a compact yet effective subset of estimators based on error correlation. Experiments on a diverse dataset of speech, singing, and music show that our proposed method with alignment outperforms individual state-of-the-art estimators in clean conditions and maintains robust voiced/unvoiced detection in noisy environments.

音高估计投票法语音处理抗噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。