arXiv:2505.12192cs.LGcs.SD2025-05被引 8

首个孟加拉语语音数据集助力帕金森病早期精准检测

BenSParX: A Robust Explainable Machine Learning Framework for Parkinson's Disease Detection from Bengali Conversational Speech

  • 构建首个孟加拉语对话语音数据集,融合多类声学特征与优化算法
  • 模型准确率达95.77%,在跨语言验证中性能超越现有方法
  • 引入SHAP分析实现预测可解释性,适合医疗AI可信赖研究

帕金森病(PD)已成为全球重大健康挑战,孟加拉国相关死亡率显著上升。在资源受限地区,基于语音的无创、低成本筛查具有潜力,但现有研究多集中于英语等主流语言,缺乏孟加拉语语音数据集,阻碍了文化包容性医疗解决方案的发展。以往研究通常仅使用有限声学特征,缺乏超参数调优与特征选择,且忽视模型可解释性,限制了模型鲁棒性与泛化能力。为此,本文提出首个面向帕金森病检测的孟加拉语对话语音数据集BenSparX,以及一个结合多样声学特征、系统特征选择与先进机器学习算法的鲁棒可解释框架。通过广泛超参数优化,并引入SHAP分析量化各特征贡献,该框架达到95.77%准确率、95.57% F1分数和0.982 AUC-ROC。进一步外部验证表明,其在其他语言数据集上仍持续优于现有方法。数据集已公开于https://github.com/Riad071/BenSParX,以促进后续研究与复现。

原文摘要 · Abstract (English)

Parkinson's disease (PD) poses a growing global health challenge, with Bangladesh experiencing a notable rise in PD-related mortality. Early detection of PD remains particularly challenging in resource-constrained settings, where voice-based analysis has emerged as a promising non-invasive and cost-effective alternative. However, existing studies predominantly focus on English or other major languages; notably, no voice dataset for PD exists for Bengali - posing a significant barrier to culturally inclusive and accessible healthcare solutions. Moreover, most prior studies employed only a narrow set of acoustic features, with limited or no hyperparameter tuning and feature selection strategies, and little attention to model explainability. This restricts the development of a robust and generalizable machine learning model. To address this gap, we present BenSparX, the first Bengali conversational speech dataset for PD detection, along with a robust and explainable machine learning framework tailored for early diagnosis. The proposed framework incorporates diverse acoustic feature categories, systematic feature selection methods, and state-of-the-art machine learning algorithms with extensive hyperparameter optimization. Furthermore, to enhance interpretability and trust in model predictions, the framework incorporates SHAP (SHapley Additive exPlanations) analysis to quantify the contribution of individual acoustic features toward PD detection. Our framework achieves state-of-the-art performance, yielding an accuracy of 95.77%, F1 score of 95.57%, and AUC-ROC of 0.982. We further externally validated our approach by applying the framework to existing PD datasets in other languages, where it consistently outperforms state-of-the-art approaches. To facilitate further research and reproducibility, the dataset has been made publicly available at https://github.com/Riad071/BenSParX.

帕金森病语音识别可解释AI孟加拉语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。