用大模型分析自发语音,识别阿尔茨海默病早期迹象。
Benchmarking Foundation Speech and Language Models for Alzheimer's Disease and Related Dementia Detection from Spontaneous Speech
- 用语音和文本大模型提取发音与语言特征
- 语音模型最高准确率73.1%,优于传统方法
- 停顿等非语义特征提升诊断效果,适合临床筛查
阿尔茨海默病及相关痴呆(ADRD)是进行性神经退行性疾病,早期检测对及时干预至关重要。自发语音包含丰富的声学与语言标记,可作为认知衰退的无创生物标志物。基础模型在大规模音频或文本数据上预训练,生成编码上下文与声学特征的高维嵌入。本研究使用PREPARE挑战数据集,包含超过1600名受试者,分为健康对照(HC)、轻度认知障碍(MCI)和阿尔茨海默病(AD)三类。排除非英语、非自发或质量差的录音后,最终数据集含703例(59.13%)HC、81例(6.81%)MCI、405例(34.06%)AD。我们基准测试了多种开源语音与语言基础模型,以分类认知状态。结果表明,Whisper-medium模型在语音模型中表现最佳(准确率=0.731,AUC=0.802);BERT结合停顿标注在文本模型中表现最优(准确率=0.662,AUC=0.744)。使用先进自动语音识别(ASR)模型生成的音频嵌入,在检测中表现优于其他方法。包含停顿等非语义特征显著提升基于文本的分类性能。结论:本研究建立了一个基于基础模型与临床相关数据集的基准框架。声学方法——尤其是基于ASR的嵌入——展现出在可扩展、无创、低成本早期检测中强大潜力。
原文摘要 · Abstract (English)
Background: Alzheimer's disease and related dementias (ADRD) are progressive neurodegenerative conditions where early detection is vital for timely intervention and care. Spontaneous speech contains rich acoustic and linguistic markers that may serve as non-invasive biomarkers for cognitive decline. Foundation models, pre-trained on large-scale audio or text data, produce high-dimensional embeddings encoding contextual and acoustic features. Methods: We used the PREPARE Challenge dataset, which includes audio recordings from over 1,600 participants with three cognitive statuses: healthy control (HC), mild cognitive impairment (MCI), and Alzheimer's Disease (AD). We excluded non-English, non-spontaneous, or poor-quality recordings. The final dataset included 703 (59.13%) HC, 81 (6.81%) MCI, and 405 (34.06%) AD cases. We benchmarked a range of open-source foundation speech and language models to classify cognitive status into the three categories. Results: The Whisper-medium model achieved the highest performance among speech models (accuracy = 0.731, AUC = 0.802). Among language models, BERT with pause annotation performed best (accuracy = 0.662, AUC = 0.744). ADRD detection using state-of-the-art automatic speech recognition (ASR) model-generated audio embeddings outperformed others. Including non-semantic features like pause patterns consistently improved text-based classification. Conclusion: This study introduces a benchmarking framework using foundation models and a clinically relevant dataset. Acoustic-based approaches -- particularly ASR-derived embeddings -- demonstrate strong potential for scalable, non-invasive, and cost-effective early detection of ADRD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。