用低秩微调大模型融合多维度语音特征,提升阿尔茨海默病早期检测精度
LoRA-Tuned Large Language Models for Dementia Detection via Multi-View Speech-Derived Features

- 通过统一提示编码四种语音特征,让大模型进行结构化多视角推理
- 在ADReSSo数据集上达到90.14%的F1分数,显著优于单一模态方法
- 无需专用编码器或后期融合,适合临床非侵入性认知筛查应用
阿尔茨海默病的早期检测有助于及时干预。自发性语言能反映认知损伤,是一种无创筛查方式。传统方法通常只关注单一表征维度——如声学特征、停顿建模、自动语音识别(ASR)转录文本或跨模态融合——限制了对异质认知症状的综合推理。本文提出一种基于低秩适应(LoRA)的大型语言模型(LLM),对四种互补的语音衍生信号进行结构化多视角推理:带停顿标记的ASR转录文本、话语级话题线索、时间流利度统计量和音位序列。这些线索被统一编码于提示中,使单一LLM可在不使用模态专用编码器或晚期融合的情况下学习一致决策函数。在ADReSSo数据集上,最佳模型达到90.14%的F1分数,消融实验验证了各视图的互补贡献。
原文摘要 · Abstract (English)
Early detection of dementia enables timely intervention, and reflecting cognitive impairment, spontaneous speech offers a non-invasive screening modality. Conventional approaches often focus on a single representational dimension -- such as acoustic descriptors, pause modeling, automatic speech recognition (ASR) transcripts, or multimodal fusion -- limiting integrative reasoning across heterogeneous cognitive symptoms. We propose a low-rank adaptation (LoRA)-tuned large language model (LLM) that performs structured multi-view reasoning over four complementary speech-derived signals: ASR transcripts with pause markers, discourse-level topic cues, temporal fluency statistics, and phonological sequences. These cues are encoded within a unified prompt, enabling a single LLM to learn a coherent decision function without modality-specific encoders or late-stage fusion. On ADReSSo, our best model achieves an F1-score of 90.14%, and ablation confirms the complementary contribution of each view.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。