arXiv:2606.28445cs.SDcs.AI2026-06中稿 · INTERSPEECH 2026

用低秩微调大模型融合多维度语音特征,提升阿尔茨海默病早期检测精度

LoRA-Tuned Large Language Models for Dementia Detection via Multi-View Speech-Derived Features

论文配图:LoRA-Tuned Large Language Models for Dementia Detection via Multi-View Speech-Derived Features
图 1 · 摘自论文原文
  • 通过统一提示编码四种语音特征,让大模型进行结构化多视角推理
  • 在ADReSSo数据集上达到90.14%的F1分数,显著优于单一模态方法
  • 无需专用编码器或后期融合,适合临床非侵入性认知筛查应用

阿尔茨海默病的早期检测有助于及时干预。自发性语言能反映认知损伤,是一种无创筛查方式。传统方法通常只关注单一表征维度——如声学特征、停顿建模、自动语音识别(ASR)转录文本或跨模态融合——限制了对异质认知症状的综合推理。本文提出一种基于低秩适应(LoRA)的大型语言模型(LLM),对四种互补的语音衍生信号进行结构化多视角推理:带停顿标记的ASR转录文本、话语级话题线索、时间流利度统计量和音位序列。这些线索被统一编码于提示中,使单一LLM可在不使用模态专用编码器或晚期融合的情况下学习一致决策函数。在ADReSSo数据集上,最佳模型达到90.14%的F1分数,消融实验验证了各视图的互补贡献。

原文摘要 · Abstract (English)

Early detection of dementia enables timely intervention, and reflecting cognitive impairment, spontaneous speech offers a non-invasive screening modality. Conventional approaches often focus on a single representational dimension -- such as acoustic descriptors, pause modeling, automatic speech recognition (ASR) transcripts, or multimodal fusion -- limiting integrative reasoning across heterogeneous cognitive symptoms. We propose a low-rank adaptation (LoRA)-tuned large language model (LLM) that performs structured multi-view reasoning over four complementary speech-derived signals: ASR transcripts with pause markers, discourse-level topic cues, temporal fluency statistics, and phonological sequences. These cues are encoded within a unified prompt, enabling a single LLM to learn a coherent decision function without modality-specific encoders or late-stage fusion. On ADReSSo, our best model achieves an F1-score of 90.14%, and ablation confirms the complementary contribution of each view.

dementia检测语音分析大模型LoRA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。