比较五种音高估计算法在次谐波语音中的表现,发现深度学习模型最准。
Comparison of fundamental frequency estimators with subharmonic voice signals
- 用语音质量分类和次谐波比(SHR)评估算法对次谐波的识别能力。
- FCN-F0在准确率和次谐波解析上均最优,优于传统方法。
- 适合临床语音分析人员选择可靠音高估计工具。
在临床语音信号分析中,次谐波发声若处理不当可能导致声学参数误判为阴性结果。因此,音高估计算法正确识别基频的能力至关重要。本文通过持续元音数据集研究了五种估计算法:Praat、YAAPT、Harvest、CREPE 和 FCN-F0。采用估计质量分类来识别次谐波错误,并使用次谐波与谐波比(SHR)衡量次谐波发声强度。结果表明,基于深度学习的 FCN-F0 在整体准确率和次谐波信号解析能力上表现最佳;CREPE 与 Harvest 也展现出优异性能,适用于持续元音分析。
原文摘要 · Abstract (English)
In clinical voice signal analysis, mishandling of subharmonic voicing may cause an acoustic parameter to signal false negatives. As such, the ability of a fundamental frequency estimator to identify speaking fundamental frequency is critical. This paper presents a sustained-vowel study, which used a quality-of-estimate classification to identify subharmonic errors and subharmonics-to-harmonics ratio (SHR) to measure the strength of subharmonic voicing. Five estimators were studied with a sustained vowel dataset: Praat, YAAPT, Harvest, CREPE, and FCN-F0. FCN-F0, a deep-learning model, performed the best both in overall accuracy and in correctly resolving subharmonic signals. CREPE and Harvest are also highly capable estimators for sustained vowel analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。