用可解释方法提升语音抑郁检测的准确率与公平性
A Fair and Transparent Framework for Speech-Based Depression Detection: Balancing Interpretability and Performance
- 结合低复杂度模型与可理解声学特征,提升模型泛化能力
- 在DAIC-WOZ数据集上实现82%测试准确率,达当前最优
- 通过可解释性分析保障临床可信度,适合医疗场景应用
语音蕴含丰富的心理健康生物标志物,但临床应用受限于模型不透明和潜在的人群偏差。本文提出一种方法框架,基于扩展版DAIC-WOZ数据集,采用随机森林、SVM和MLP等低复杂度机器学习模型,结合可理解的声学特征(MFCCs, eGeMAPS),以减轻过拟合并增强泛化性。为平衡准确性与临床信任,引入LIME和SHAP等可解释性方法进行特征选择,并通过统计显著性检验和人群公平性分析,消除虚假相关性。实验表明,经XAI优化的特征子集与MLP结合,在测试集上达到82%的准确率,优于现有方法。本研究构建了一个透明、鲁棒且符合伦理的辅助技术框架,可推广至其他二分类任务。
原文摘要 · Abstract (English)
While speech provides rich, non-invasive biomarkers for mental-health assessment, clinical adoption is limited by opaque models and potential demographic bias. In this work we propose a methodological framework to evaluate robustness and interpretability for automated depression detection on the extended DAIC-WOZ dataset using low-complexity machine learning baselines (RF, SVM, and MLP) chosen to mitigate overfitting and enhance generalization in combination with human-understandable acoustic features (MFCCs, eGeMAPS). To balance accuracy with clinical trust, we leverage explainability methods (LIME and SHAP) for feature selection, validating our findings with statistical significance tests and demographic fairness analyses to mitigate spurious, artifact-driven correlations. Empirical results demonstrate that an optimized subset of explainable AI (XAI)-selected features combined with an MLP architecture achieves a state-of-the-art test accuracy of 82\%. Ultimately, this work provides a transparent framework for robust and ethical assistive technologies that can be applied to any other binary task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。