用声音数据自动识别嗓音疾病,无需侵入检查。
Voice Pathology Detection Using Phonation
- 基于声学特征和循环神经网络分析语音数据
- 在萨尔布吕肯数据库上达到92.3%分类准确率
- 适合医疗AI、语音健康监测场景使用
嗓音障碍严重影响沟通与生活质量,亟需早期精准诊断。传统喉镜检查具有侵入性、主观性强且难以获取。本研究提出一种基于机器学习的非侵入式嗓音病理检测框架,利用萨尔布吕肯语音数据库中的发声数据,提取梅尔频率倒谱系数(MFCCs)、音高特征与梅尔频谱图等声学特征。采用包含LSTM与注意力机制的循环神经网络对样本进行正常与病理性分类。通过音高偏移与高斯噪声添加等数据增强技术提升模型泛化能力,预处理保障信号质量。进一步引入霍尔德指数与赫斯特指数等尺度特征,捕捉信号不规则性与长期依赖关系。该框架提供了一种非侵入、自动化嗓音病理早期检测工具,助力AI医疗发展,改善患者预后。
原文摘要 · Abstract (English)
Voice disorders significantly affect communication and quality of life, requiring an early and accurate diagnosis. Traditional methods like laryngoscopy are invasive, subjective, and often inaccessible. This research proposes a noninvasive, machine learning-based framework for detecting voice pathologies using phonation data. Phonation data from the Saarbrücken Voice Database are analyzed using acoustic features such as Mel Frequency Cepstral Coefficients (MFCCs), chroma features, and Mel spectrograms. Recurrent Neural Networks (RNNs), including LSTM and attention mechanisms, classify samples into normal and pathological categories. Data augmentation techniques, including pitch shifting and Gaussian noise addition, enhance model generalizability, while preprocessing ensures signal quality. Scale-based features, such as Hölder and Hurst exponents, further capture signal irregularities and long-term dependencies. The proposed framework offers a noninvasive, automated diagnostic tool for early detection of voice pathologies, supporting AI-driven healthcare, and improving patient outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。