arXiv:2505.18722eess.AScs.AI2025-05中稿 · Interspeech 2025被引 3

用非诊断语音数据也能准确识别帕金森病,效果不输专业数据集。

Evaluating the Usefulness of Non-Diagnostic Speech Data for Developing Parkinson's Disease Classifiers

  • 利用非诊断性语音数据(如对话轮替)进行帕金森病分类
  • 非诊断数据与专业数据集表现相当,拼接音频+平衡性别和状态提升效果
  • 模型在非诊断数据上训练更具泛化能力,适合实际应用

基于语音的帕金森病(PD)检测因其自动化、低成本和非侵入性而受到关注。现有研究多依赖于诊断导向的语音任务数据,本文探索了使用非诊断目的语音数据(如Turn-Taking, TT数据集)进行PD分类的可行性。结果表明,TT数据集在性能上可媲美专业诊断数据集PC-GITA。研究还发现,拼接音频片段以及平衡参与者性别与疾病状态分布有助于提升分类效果。跨数据集评估显示,以PC-GITA训练的模型在TT上泛化能力差,而以TT训练的模型在PC-GITA上表现更优。此外,分析揭示了不同交叉验证折间存在较大变异性,主要源于个体说话人表现差异显著。

原文摘要 · Abstract (English)

Speech-based Parkinson's disease (PD) detection has gained attention for its automated, cost-effective, and non-intrusive nature. As research studies usually rely on data from diagnostic-oriented speech tasks, this work explores the feasibility of diagnosing PD on the basis of speech data not originally intended for diagnostic purposes, using the Turn-Taking (TT) dataset. Our findings indicate that TT can be as useful as diagnostic-oriented PD datasets like PC-GITA. We also investigate which specific dataset characteristics impact PD classification performance. The results show that concatenating audio recordings and balancing participants' gender and status distributions can be beneficial. Cross-dataset evaluation reveals that models trained on PC-GITA generalize poorly to TT, whereas models trained on TT perform better on PC-GITA. Furthermore, we provide insights into the high variability across folds, which is mainly due to large differences in individual speaker performance.

帕金森病语音识别非诊断数据分类器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。