分析语音模型在认知障碍检测中的偏见,发现女性和年轻人易误诊。
Bias and Fairness in Self-Supervised Acoustic Representations for Cognitive Impairment Detection
- 对比传统特征与Wav2Vec 2.0嵌入,评估跨群体分类表现
- 女性和年轻群体的识别准确率低至AUC 0.746,特异性差距达18%
- 揭示模型在真实临床场景中存在系统性偏差,需针对性优化
基于语音的认知障碍(CI)检测提供了早期非侵入式诊断的可能,但不同人口与临床子群体间的表现差异尚未充分探讨,引发公平性与泛化性的担忧。本研究对DementiaBank Pitt语料库中的语音特征进行了系统的偏见分析,比较了传统声学特征(MFCCs, eGeMAPS)与Wav2Vec 2.0(W2V2)的上下文嵌入在CI与抑郁分类中的表现。对于CI检测,高层W2V2嵌入优于基线特征(平均准确率最高达80.6%),但表现出显著性能差异:女性和年轻参与者判别能力较弱(AUC分别为0.769和0.746),特异性差距最大分别达18%和15%,导致误诊风险更高。这些差异反映代表偏差——即模型在不同人口或临床群体中表现的系统性差异。在CI患者中进行抑郁检测的整体性能较低,仅低层与中层W2V2嵌入带来轻微提升。跨任务泛化能力有限,表明两任务依赖不同的表征。研究强调,在临床语音应用中需开展公平性感知的模型评估与子群体分析,尤其面对现实世界中的异质性人群。
原文摘要 · Abstract (English)
Speech-based detection of cognitive impairment (CI) offers a promising non-invasive approach for early diagnosis, yet performance disparities across demographic and clinical subgroups remain underexplored, raising concerns around fairness and generalizability. This study presents a systematic bias analysis of acoustic-based CI and depression classification using the DementiaBank Pitt Corpus. We compare traditional acoustic features (MFCCs, eGeMAPS) with contextualized speech embeddings from Wav2Vec 2.0 (W2V2), and evaluate classification performance across gender, age, and depression-status subgroups. For CI detection, higher-layer W2V2 embeddings outperform baseline features (UAR up to 80.6\%), but exhibit performance disparities; specifically, females and younger participants demonstrate lower discriminative power (\(AUC\): 0.769 and 0.746, respectively) and substantial specificity disparities (\(Δ_{spec}\) up to 18\% and 15\%, respectively), leading to a higher risk of misclassifications than their counterparts. These disparities reflect representational biases, defined as systematic differences in model performance across demographic or clinical subgroups. Depression detection within CI subjects yields lower overall performance, with mild improvements from low and mid-level W2V2 layers. Cross-task generalization between CI and depression classification is limited, indicating that each task depends on distinct representations. These findings emphasize the need for fairness-aware model evaluation and subgroup-specific analysis in clinical speech applications, particularly in light of demographic and clinical heterogeneity in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。