跨语言阿尔茨海默病语音检测,用多模态融合提升零样本迁移能力
Synergizing Zero-Shot Cross-Lingual Alzheimer Detection with Language-Invariant Multimodal Bi-Geometric Adversarial Learning

- 融合多语言语音与文本模型,学习跨语言通用表征
- 在零样本跨语言测试中,多模态性能显著优于单模态
- 适合做跨语言医疗语音分析与阿尔茨海默病早期筛查
本文研究零样本跨语言语音基阿尔茨海默病检测(SADD)。我们假设,通过融合多语言语音与文本预训练模型来学习语言无关的多模态表征,对可靠迁移到未见语言至关重要,因为两种模态分别捕捉认知障碍的声学与语言特征,而对抗学习可抑制语言特异性干扰。零样本跨语言评估结果验证了该假设,表明多模态融合始终优于单模态基线。为此,我们提出 ORBIT 框架,结合交叉注意力融合、多触点语言对抗器、互补球面-双曲几何学习与共识聚类。在各类设置下,ORBIT 性能均强于单模态模型和简单的拼接融合基线。
原文摘要 · Abstract (English)
In this work, we study zero-shot cross-lingual speech-based Alzheimer's disease detection (SADD). We hypothesize that learning language-invariant multimodal representations by fusing multilingual speech and text pretrained models is essential for reliable transfer to unseen languages, as the two modalities capture complementary acoustic and linguistic markers of cognitive impairment while adversarial learning suppresses language-specific confounds. Empirical results in zero-shot cross-lingual evaluation substantiate the hypothesis, showing that multimodal fusion consistently outperforms unimodal baselines. To this end, we propose ORBIT, a novel framework that combines cross-attentive fusion, multi-tap language adversaries, and complementary spherical--hyperbolic geometric learning with consensus clustering. Across settings, ORBIT achieves the strongest performance compared to unimodal models and simple concatenation-based fusion baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。