AI检测认知障碍时对多语者有偏见,可能误判健康人群。
Can we trust AI to detect healthy multilingual English speakers among the cognitively impaired cohort in the UK? An investigation using real-world conversational speech
- 用语音特征分析多语者,发现模型存在系统性偏差。
- 多语者被误判为认知衰退的概率更高,尤其在南约克方言中。
- 当前模型不适用于少数族裔多语者,需改进以减少偏见。
对话语音常能揭示认知衰退的早期迹象,如痴呆和轻度认知障碍。在英国,每四人中就有一人属于少数族裔,而黑人与亚裔群体的痴呆发病率预计将上升最快。本研究考察了人工智能模型在识别认知受损群体中健康多语英语使用者时的可信度,重点检验其是否存在偏见,以确保临床适用性。实验招募了全国范围内的单语者,以及谢菲尔德和布拉德福德四个社区中心的多语者,涵盖索马里语、汉语及南亚语言使用者,并进一步按西约克与南约克口音划分,以严格测试模型性能。尽管自动语音识别(ASR)系统未显示组间显著偏差,但基于声学与语言特征的分类与回归模型对多语者表现出明显偏见,尤其在记忆、流利度和阅读任务中。该偏见在使用公开数据集DementiaBank训练的模型中更为突出,且多语者更易被错误归类为存在认知衰退。本研究首次发现,尽管整体表现良好,现有AI模型仍对英国少数族裔背景的多语者存在偏见,且特定口音(南约克)使用者更可能被误判为病情更重。本试点研究结论:当前AI工具尚不可靠用于此类人群诊断,未来将致力于开发更具泛化能力、减缓偏见的模型。
原文摘要 · Abstract (English)
Conversational speech often reveals early signs of cognitive decline, such as dementia and MCI. In the UK, one in four people belongs to an ethnic minority, and dementia prevalence is expected to rise most rapidly among Black and Asian communities. This study examines the trustworthiness of AI models, specifically the presence of bias, in detecting healthy multilingual English speakers among the cognitively impaired cohort, to make these tools clinically beneficial. For experiments, monolingual participants were recruited nationally (UK), and multilingual speakers were enrolled from four community centres in Sheffield and Bradford. In addition to a non-native English accent, multilinguals spoke Somali, Chinese, or South Asian languages, who were further divided into two Yorkshire accents (West and South) to challenge the efficiency of the AI tools thoroughly. Although ASR systems showed no significant bias across groups, classification and regression models using acoustic and linguistic features exhibited bias against multilingual speakers, particularly in memory, fluency, and reading tasks. This bias was more pronounced when models were trained on the publicly available DementiaBank dataset. Moreover, multilinguals were more likely to be misclassified as having cognitive decline. This study is the first of its kind to discover that, despite their strong overall performance, current AI models show bias against multilingual individuals from ethnic minority backgrounds in the UK, and they are also more likely to misclassify speakers with a certain accent (South Yorkshire) as living with a more severe cognitive decline. In this pilot study, we conclude that the existing AI tools are therefore not yet reliable for diagnostic use in these populations, and we aim to address this in future work by developing more generalisable, bias-mitigated models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。