arXiv:2506.02078eess.AScs.AI2025-06中稿 · Interspeech 2025被引 2

用预训练音频嵌入识别帕金森语音,OpenL3表现最佳

Evaluating the Effectiveness of Pre-Trained Audio Embeddings for Classification of Parkinson's Disease Speech Data

  • 对比三种预训练音频模型在语音中提取特征
  • OpenL3在发音测试中准确率最高,尤其擅长捕捉关键语音特征
  • 仅Wav2Vec2.0存在性别偏差,适合关注模型公平性的研究者

帕金森病(PD)的语音障碍是常见生物标志物,推动了基于语音数据的诊断技术发展。尽管深度声学特征在PD分类中展现出潜力,但其效果常受个体差异影响,这一问题在现有研究中尚未充分探讨。本研究评估了三种预训练音频嵌入模型(OpenL3、VGGish 和 Wav2Vec2.0)在帕金森病语音分类中的表现。基于NeuroVoz数据集,OpenL3在双音节快速交替发音(DDK)和听读任务(LR)中优于其他模型,有效捕捉了用于PD检测的关键声学特征。仅Wav2Vec2.0在DDK任务中表现出显著性别偏差,对男性说话者结果更优。误分类案例揭示了异常语音模式带来的挑战,强调了提升特征提取能力和模型鲁棒性的必要性。

原文摘要 · Abstract (English)

Speech impairments are prevalent biomarkers for Parkinson's Disease (PD), motivating the development of diagnostic techniques using speech data for clinical applications. Although deep acoustic features have shown promise for PD classification, their effectiveness often varies due to individual speaker differences, a factor that has not been thoroughly explored in the existing literature. This study investigates the effectiveness of three pre-trained audio embeddings (OpenL3, VGGish and Wav2Vec2.0 models) for PD classification. Using the NeuroVoz dataset, OpenL3 outperforms others in diadochokinesis (DDK) and listen and repeat (LR) tasks, capturing critical acoustic features for PD detection. Only Wav2Vec2.0 shows significant gender bias, achieving more favorable results for male speakers, in DDK tasks. The misclassified cases reveal challenges with atypical speech patterns, highlighting the need for improved feature extraction and model robustness in PD detection.

语音识别帕金森病预训练模型医学应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。