用多模态大模型统一语音与文本,提升阿尔茨海默病和帕金森病的分期准确率。
Unifying Acoustic Features and Text with Multimodal LLMs for Neurodegenerative Screening

- 将音频频谱图和文本通过视觉变换器编码,融合进大语言模型嵌入空间。
- 在Bridge2AI-Voice数据集上,分期准确率超越传统方法和现有LLM方案。
- 适合医疗筛查场景,支持轻量化部署,可扩展至多种神经退行性疾病。
基于语音的筛查为评估阿尔茨海默病(AD)和帕金森病(PD)等神经退行性疾病提供了可扩展且无创的途径,但因异构数据难以整合,疾病分期仍具挑战。本文提出NeurMLLM,一种高效多模态生成框架,用于神经退行性疾病分期。NeurMLLM首先使用视觉变换器对音频的频谱图和梅尔频率倒谱系数进行编码,并将其表示投影到大语言模型(LLM)的嵌入空间中,再与转录文本及人口统计学指令标记拼接为单一统一序列。随后,通过低秩适应(LoRA)进行指令微调,使LLM以自回归方式预测受限标签标记,实现生成式分类。在Bridge2AI-Voice数据集上对AD和PD进行细粒度分期评估,结果表明NeurMLLM表现优异,持续优于经典机器学习方法和现有基于LLM的方法。研究表明,多模态大模型在神经退行性疾病分期中具有巨大潜力,不仅能提升分期准确性,还支持便捷部署。
原文摘要 · Abstract (English)
Voice-based screening offers a scalable and non-invasive way to assess neurodegenerative diseases such as Alzheimer's disease (AD) and Parkinson's disease (PD), but their staging remains challenging due to the difficulty of integrating heterogeneous data. This paper presents NeurMLLM, an efficient multimodal generative framework for neurodegenerative disease staging. NeurMLLM first encodes the spectrograms and Mel-frequency cepstral coefficients of audio data with vision transformers and projects their representations into the embedding space of a large language model (LLM), where they are concatenated with transcript and demographic instruction tokens as a single unified sequence. The LLM is then instruction-tuned via Low-Rank Adaptation using task prompts to autoregressively predict a constrained label token, enabling a generative classification. By evaluating on the Bridge2AI-Voice dataset for fine-grained staging of AD and PD, we observe that NeurMLLM achieves strong performance, consistently outperforming classical machine learning methods and existing LLM-based approaches. The results show the high potential of multimodal LLMs in neurodegenerative disease staging, improving staging accuracy and supporting accessible deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。