arXiv:2602.14655cs.CLcs.AI2026-02中稿 · ICASSP 2026 confer…被引 1

用联邦学习+语音增强,解决阿尔茨海默病诊断数据少且难共享的难题

Breaking Data Efficiency Dilemma: A Federated and Augmented Learning Framework For Alzheimer's Disease Detection via Speech

  • 通过跨类别语音内容重组生成病理语音样本,提升数据多样性
  • 在隐私保护下实现多机构协作,模型准确率达91.52%创纪录
  • 适合医疗AI研究者和需要隐私安全训练的团队使用

阿尔茨海默病早期诊断对延缓病情进展至关重要。尽管基于人工智能的语音检测具有无创、低成本优势,但受限于医疗数据稀缺与隐私保护壁垒,面临严重数据效率困境。为此,我们提出FAL-AD框架,融合联邦学习与数据增强,系统性优化数据效率。该方法实现三大突破:首先,基于语音转换的数据增强生成多样病理语音样本;其次,采用自适应联邦学习范式,在隐私约束下实现跨机构协同优化;最后,通过注意力驱动的跨模态融合模型,实现词级对齐与声学-文本交互。在ADReSSo数据集上,FAL-AD达到91.52%的多模态分类准确率,超越所有集中式基线,为数据效率难题提供实用解决方案。源代码已公开于https://github.com/smileix/fal-ad。

原文摘要 · Abstract (English)

Early diagnosis of Alzheimer's Disease (AD) is crucial for delaying its progression. While AI-based speech detection is non-invasive and cost-effective, it faces a critical data efficiency dilemma due to medical data scarcity and privacy barriers. Therefore, we propose FAL-AD, a novel framework that synergistically integrates federated learning with data augmentation to systematically optimize data efficiency. Our approach delivers three key breakthroughs: First, absolute efficiency improvement through voice conversion-based augmentation, which generates diverse pathological speech samples via cross-category voice-content recombination. Second, collaborative efficiency breakthrough via an adaptive federated learning paradigm, maximizing cross-institutional benefits under privacy constraints. Finally, representational efficiency optimization by an attentive cross-modal fusion model, which achieves fine-grained word-level alignment and acoustic-textual interaction. Evaluated on ADReSSo, FAL-AD achieves a state-of-the-art multi-modal accuracy of 91.52%, outperforming all centralized baselines and demonstrating a practical solution to the data efficiency dilemma. Our source code is publicly available at https://github.com/smileix/fal-ad.

阿尔茨海默病联邦学习语音分析数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。