通过消除性别语言偏见,提升多模态认知障碍检测公平性
Fair Cognitive Impairment Detection Through Unlearning

- 融合语音文本图像多模态信息,用梯度反向抑制无关人口属性
- 在跨语言数据集上性能超越现有方法,男女/语言组间差距缩小40%以上
- 适合关注医疗AI公平性的研究者与临床筛查系统开发者
轻度认知障碍(MCI)表现为记忆力、语言或思维能力的明显下降。从自发性言语中检测MCI具有大规模筛查的潜力。然而,现有模型常利用与标签相关的年龄、性别等人口统计学线索,导致不同子群体间性能差异显著。本文提出一种多模态框架,结合(i)语音、文本和图像之间的跨模态融合,以及(ii)基于梯度反向的去偏学习,使共享嵌入不编码与任务无关的人口属性。在多语言基准TAUKADIAL和PREPARE上评估,该方法在MCI分类任务中优于当前最优的多语言及多模态基线模型,同时显著降低患者子群体(性别与语言)间的性能差距。进一步分析跨数据集迁移发现,去偏学习有助于构建更鲁棒的MCI检测表示。
原文摘要 · Abstract (English)
Mild Cognitive Impairment (MCI) is a medical condition characterized by a noticeable decline in memory, language, or thinking abilities. MCI detection from spontaneous speech is promising for scalable screening. However, learned models often exploit demographic cues correlated with labels, resulting in a large performance gap across subgroups. We present a multimodal framework that combines (i) cross-model fusion between modalities (speech, text, and image), and (ii) unlearning using gradient reversal that discourages the shared embedding from encoding task-irrelevant demographic attributes. Evaluated on the multilingual benchmarks TAUKADIAL and PREPARE, our method outperforms the state-of-the-art multilingual and multimodal baseline in MCI classification while substantially reducing the performance gap across patient subgroups (sex and language). We further analyze transfer across datasets, showing that demographic unlearning helps learn more robust representations for MCI detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。