arXiv:2601.16989eess.AScs.CL2026-01被引 2

首次系统评估语音认知障碍检测中的公平性,发现模型架构与人群特征需共同考虑。

The Voice of Equity: A Systematic Evaluation of Bias Mitigation Techniques for Speech-Based Cognitive Impairment Detection Across Architectures and Demographics

  • 构建双模型并对比预处理、内处理、后处理多种公平性干预方法
  • 80岁以上群体敏感性低,西语使用者真阳性率显著偏低
  • 不同架构对同一干预效果差异大,需个性化设计公平策略

基于语音的认知障碍检测具备可扩展、无创筛查潜力,但算法在不同人口与语言群体中的偏差问题尚未充分研究。本文提出首个针对多类别认知障碍语音检测的系统性公平性分析框架,全面评估了不同架构与人口子群体中的偏差缓解效果。在多语言NIA PREPARE Challenge数据集上,我们开发了两个基于Transformer的模型:SpeechCARE-AGF(F1: 70.87)和Whisper-LWF-LoRA(F1: 71.46)。相较于以往仅考察单一缓解方法的研究,本工作系统比较了预处理、内处理与后处理策略,采用平等机会与均等化奇偶性作为公平性度量,覆盖性别、年龄、教育水平与语言等维度。结果发现,80岁及以上人群敏感性较低,西班牙语使用者真阳性率(TPR)显著低于英语使用者。不同架构对干预响应差异明显:过采样使SpeechCARE-AGF在80+群体中TPR从46.19%提升至49.97%,但对Whisper-LWF-LoRA影响极小。研究揭示模型架构从根本上决定偏差模式与缓解有效性,自适应融合机制支持灵活应对数据干预,频段重加权则在两类模型中均带来稳健改进。结论强调,公平性干预必须兼顾模型架构与人口特征,为构建公平的语音筛查工具提供系统性框架。

原文摘要 · Abstract (English)

Speech-based detection of cognitive impairment offers a scalable, non-invasive screening, yet algorithmic bias across demographic and linguistic subgroups remains critically underexplored. We present the first comprehensive fairness analysis framework for speech-based multi-class cognitive impairment detection, systematically evaluating bias mitigation across architectures, and demographic subgroups. We developed two transformer-based architectures, SpeechCARE-AGF and Whisper-LWF-LoRA, on the multilingual NIA PREPARE Challenge dataset. Unlike prior work that typically examines single mitigation techniques, we compared pre-processing, in-processing, and post-processing approaches, assessing fairness via Equality of Opportunity and Equalized Odds across gender, age, education, and language. Both models achieved strong performance (F1: SpeechCARE-AGF 70.87, Whisper-LWF-LoRA 71.46) but exhibited substantial fairness disparities. Adults >=80 showed lower sensitivity versus younger groups; Spanish speakers demonstrated reduced TPR versus English speakers. Mitigation effectiveness varied by architecture: oversampling improved SpeechCARE-AGF for older adults (80+ TPR: 46.19%=>49.97%) but minimally affected Whisper-LWF-LoRA. This study addresses a critical healthcare AI gap by demonstrating that architectural design fundamentally shapes bias patterns and mitigation effectiveness. Adaptive fusion mechanisms enable flexible responses to data interventions, while frequency reweighting offers robust improvements across architectures. Our findings establish that fairness interventions must be tailored to both model architecture and demographic characteristics, providing a systematic framework for developing equitable speech-based screening tools essential for reducing diagnostic disparities in cognitive healthcare.

语音识别公平性认知障碍医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。