用大模型生成语音数据,提升阿尔茨海默病早期筛查准确率
LLMCARE: early detection of cognitive impairment via transformer models enhanced by LLM-generated synthetic data
- 融合变压器模型与语言特征,构建多模态筛查框架
- 大模型合成数据使准确率提升至F1=85.7(ADReSSo数据集)
- 适合关注老年认知障碍早筛的临床与算法研究者
阿尔茨海默病及相关痴呆症影响美国近五百万老年人,但超过一半未被诊断。基于语音的自然语言处理提供了一种可扩展的早期认知衰退检测方法,通过细微语言特征在临床诊断前发现异常。本研究开发并评估了一个整合变压器嵌入、手工语言特征、大语言模型(LLM)生成合成数据的语音筛查流程,并对比了单模态与多模态分类器性能。外部验证在仅含轻度认知障碍(MCI)的独立队列中进行。数据来自ADReSSo 2021基准数据集(n=237,Pitt Corpus)和DementiaBank Delaware语料库(n=205,MCI vs. 对照)。测试了十种变压器模型及三种微调策略。最终采用晚期融合模型结合最优变压器嵌入与110个语言特征。五个LLM(LLaMA8B/70B、MedAlpaca7B、Ministral8B、GPT-4o)生成标签条件合成语音用于数据增强,三个多模态LLM(GPT-4o、Qwen-Omni、Phi-4)在零样本与微调模式下评估。在ADReSSo上,融合模型达到F1=83.3(AUC=89.5),优于纯变压器与语言基线;使用MedAlpaca7B(2×)增强后F1提升至85.7,但规模过大时收益递减。微调显著提升单模态模型表现(MedAlpaca7B F1从47.7升至78.7),而多模态模型性能仍较低(Phi-4=71.6;GPT-4o=67.6)。在Delaware数据集上,融合+1×MedAlpaca7B模型达F1=72.8(AUC=69.6)。结果表明,融合变压器与语言特征可有效提升ADRD检测能力;基于LLM的数据增强提升数据效率但存在边际效应,当前多模态模型仍有局限。在独立MCI队列上的验证支持该流程在可扩展、临床相关早筛中的潜力。
原文摘要 · Abstract (English)
Alzheimer's disease and related dementias(ADRD) affect nearly five million older adults in the United States, yet more than half remain undiagnosed. Speech-based natural language processing(NLP) offers a scalable approach for detecting early cognitive decline through subtle linguistic markers that may precede clinical diagnosis. This study develops and evaluates a speech-based screening pipeline integrating transformer embeddings with handcrafted linguistic features, synthetic augmentation using large language models(LLMs), and benchmarking of unimodal and multimodal classifiers. External validation assessed generalizability to a MCI-only cohort. Transcripts were drawn from the ADReSSo 2021 benchmark dataset(n=237, Pitt Corpus) and the DementiaBank Delaware corpus(n=205, MCI vs. controls). Ten transformer models were tested under three fine-tuning strategies. A late-fusion model combined embeddings from the top transformer with 110 linguistic features. Five LLMs(LLaMA8B/70B, MedAlpaca7B, Ministral8B,GPT-4o) generated label-conditioned synthetic speech for augmentation, and three multimodal LLMs(GPT-4o,Qwen-Omni,Phi-4) were evaluated in zero-shot and fine-tuned modes. On ADReSSo, the fusion model achieved F1=83.3(AUC=89.5), outperforming transformer-only and linguistic baselines. MedAlpaca7B augmentation(2x) improved F1=85.7, though larger scales reduced gains. Fine-tuning boosted unimodal LLMs(MedAlpaca7B F1=47.7=>78.7), while multimodal models performed lower (Phi-4=71.6;GPT-4o=67.6). On Delaware, the fusion plus 1x MedAlpaca7B model achieved F1=72.8(AUC=69.6). Integrating transformer and linguistic features enhances ADRD detection. LLM-based augmentation improves data efficiency but yields diminishing returns, while current multimodal models remain limited. Validation on an independent MCI cohort supports the pipeline's potential for scalable, clinically relevant early screening.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。