arXiv:2504.08997eess.AS2025-04被引 2

分析语音障碍检测系统在不同人群中的公平性,发现全局指标掩盖了性别年龄偏差。

Beyond Global Metrics: A Fairness Analysis for Interpretable Voice Disorder Detection Systems

  • 按性别年龄分组评估系统性能,发现存在系统性误判
  • 老年健康者易被误判为患病,青年患者常被漏诊
  • 针对群体的校准可减少过自信,提升可靠性

我们基于包含人口统计学元数据的现有语音障碍数据集,对自动语音障碍检测(AVDD)系统进行了全面分析。研究考察了系统在不同人口群体中的表现,重点关注性别和年龄分组。性能评估采用归一化成本和交叉熵等多指标。通过在预定义的人口组上分别训练校准技术,缓解了群体依赖的校准偏差。分析显示,尽管全局指标表现良好,但各群体间存在显著性能差异。系统呈现系统性偏差:55岁以上健康者常被误判为有语音障碍,14-30岁患者则常被误判为健康。群体特定校准改善了后验概率质量,降低了过自信。对于年轻患者,低严重度评分是导致性能不佳的原因;对于老年人,年龄相关声学特征及预训练Hubert模型作为特征提取器的局限性可能影响结果。研究表明,仅靠全局指标不足以评估AVDD系统性能。群体分析可揭示隐藏在全局指标下的问题。群体依赖校准策略有助于减轻偏差,提升系统置信度可靠性。这些发现强调了在语音障碍检测中引入人口统计学特异性评估与校准的必要性,并为其他具有人口统计学元数据的生物医学分类任务提供了方法框架。

原文摘要 · Abstract (English)

We conducted a comprehensive analysis of an Automatic Voice Disorders Detection (AVDD) system using existing voice disorder datasets with available demographic metadata. The study involved analysing system performance across various demographic groups, particularly focusing on gender and age-based cohorts. Performance evaluation was based on multiple metrics, including normalised costs and cross-entropy. We employed calibration techniques trained separately on predefined demographic groups to address group-dependent miscalibration. Analysis revealed significant performance disparities across groups despite strong global metrics. The system showed systematic biases, misclassifying healthy speakers over 55 as having a voice disorder and speakers with disorders aged 14-30 as healthy. Group-specific calibration improved posterior probability quality, reducing overconfidence. For young disordered speakers, low severity scores were identified as contributing to poor system performance. For older speakers, age-related voice characteristics and potential limitations in the pretrained Hubert model used as feature extractor likely affected results. The study demonstrates that global performance metrics are insufficient for evaluating AVDD system performance. Group-specific analysis may unmask problems in system performance which are hidden within global metrics. Further, group-dependent calibration strategies help mitigate biases, resulting in a more reliable indication of system confidence. These findings emphasize the need for demographic-specific evaluation and calibration in voice disorder detection systems, while providing a methodological framework applicable to broader biomedical classification tasks where demographic metadata is available.

语音检测公平性分析群体校准生物医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。