arXiv:2501.09159eess.AScs.SD2025-01

用卷积网络自动识别语音中的病理性次谐波,准确率超98%

Towards detecting the pathological subharmonic voicing with fully convolutional neural networks

  • 采用全卷积神经网络分析语音信号整体特征
  • 合成数据训练下分类准确率达98%以上
  • 适合嗓音疾病诊断与语音病理学研究者

许多嗓音障碍会导致次谐波发声,但目前缺乏可靠的次谐波检测技术。区分次谐波发声与正常发声极具挑战,因二者均为近周期现象。次谐波发声会在正常声门周期中引入周期性变化,因此需对信号进行整体分析以估计次谐波周期。深度学习是解决此类复杂问题的有效方法。本文提出基于合成次谐波语音信号训练的全卷积神经网络,用于分类次谐波周期。合成数据评估显示分类准确率超过98%,对持续元音录音的评估也取得令人鼓舞的结果,并指出了未来改进方向。

原文摘要 · Abstract (English)

Many voice disorders induce subharmonic phonation, but voice signal analysis is currently lacking a technique to detect the presence of subharmonics reliably. Distinguishing subharmonic phonation from normal phonation is a challenging task as both are nearly periodic phenomena. Subharmonic phonation adds cyclical variations to the normal glottal cycles. Hence, the estimation of subharmonic period requires a wholistic analysis of the signals. Deep learning is an effective solution to this type of complex problem. This paper describes fully convolutional neural networks which are trained with synthesized subharmonic voice signals to classify the subharmonic periods. Synthetic evaluation shows over 98% classification accuracy, and assessment of sustained vowel recordings demonstrates encouraging outcomes as well as the areas for future improvements.

语音病理深度学习次谐波

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。