优化大模型适配策略,让语音分析更准地筛查阿尔茨海默病。
Speech-Based Cognitive Screening: A Systematic Evaluation of LLM Adaptation Strategies
- 用演示样本选择和推理提示提升大模型性能
- 微调后文本模型表现最优,超过多数商用系统
- 适合医疗筛查场景,尤其关注低成本可扩展方案
超过一半美国阿尔茨海默病及相关痴呆(ADRD)患者未被诊断,语音筛查提供了可扩展的检测路径。本研究基于DementiaBank语音语料库,对比了九个纯文本模型与三个多模态音视频-文本模型在痴呆检测中的表现,评估了上下文学习、推理增强提示、参数高效微调及多模态融合等适配策略。结果表明:以类别中心样本为示范的上下文学习效果最佳;推理提示对小模型有提升作用;标记级微调通常取得最高得分;增加分类头显著改善低性能模型表现。多模态模型中,微调后的音视频-文本系统表现良好,但未超越顶尖纯文本模型。研究强调,模型适配策略(包括示范选择、推理设计和微调方法)对语音痴呆检测至关重要,且经过合理适配的开源模型可达到甚至超过商业系统水平。
原文摘要 · Abstract (English)
Over half of US adults with Alzheimer disease and related dementias remain undiagnosed, and speech-based screening offers a scalable detection approach. We compared large language model adaptation strategies for dementia detection using the DementiaBank speech corpus, evaluating nine text-only models and three multimodal audio-text models on recordings from DementiaBank speech corpus. Adaptations included in-context learning with different demonstration selection policies, reasoning-augmented prompting, parameter-efficient fine-tuning, and multimodal integration. Results showed that class-centroid demonstrations achieved the highest in-context learning performance, reasoning improved smaller models, and token-level fine-tuning generally produced the best scores. Adding a classification head substantially improved underperforming models. Among multimodal models, fine-tuned audio-text systems performed well but did not surpass the top text-only models. These findings highlight that model adaptation strategies, including demonstration selection, reasoning design, and tuning method, critically influence speech-based dementia detection, and that properly adapted open-weight models can match or exceed commercial systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。