arXiv:2607.21496eess.SPcs.LG2026-07

用语音多模态大模型实现更通用的认知障碍检测

Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models

论文配图:Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models
图 1 · 摘自论文原文
  • 融合语音音频与转录文本,通过大模型提取多模态特征
  • 在两个数据集上达到92.4%准确率,优于单一模态方法
  • 保护隐私且可跨数据集泛化,适合临床筛查场景

认知障碍(CI)是日益严峻的公共健康问题。早期精准诊断对及时干预和改善患者预后至关重要。基于语音的CI检测因其非侵入性而成为有前景的方法,语音信号同时包含语言和声学特征,与认知衰退相关。大语言模型(LLMs)的发展提升了语音评估的表达能力与跨说话人、设备及临床环境的泛化能力。通过联合建模语言与声学特征的多模态学习,能更全面刻画与CI相关的认知与行为变化,提升检测可靠性。本文提出一种基于开源大语言模型的多模态CI检测框架,整合语音音频与对应转录文本,同时保护患者隐私。声学嵌入直接从语音信号提取,文本嵌入由自动转写生成,两者拼接后用于下游分类,无需访问原始或敏感数据。在ADReSS20和ADReSSo21基准数据集上评估,所提方法取得92.4%的分类准确率,显著优于单模态基线。该工作建立了新的最优性能,证明了基于大模型的多模态融合在实现鲁棒、可扩展、非侵入式筛查方面的潜力。

原文摘要 · Abstract (English)

Cognitive impairment (CI) is a growing public health concern. Early and accurate diagnosis is critical for enabling timely intervention and improving patient outcomes. Speech-based CI detection has emerged as a promising non-invasive approach, as speech signals encode both linguistic and acoustic markers associated with cognitive decline. Recent advances in large language models (LLMs) further strengthen the potential of speech-based assessment by enabling more expressive representation learning and improved generalization across diverse speakers, recording devices, and clinical environments. Moreover, multimodal learning by jointly modeling linguistic and acoustic features allows for a more comprehensive characterization of cognitive and behavioral changes related to CI, leading to more reliable detection. In this work, we propose a multimodal CI detection framework based on open-source LLMs that integrates speech audio and corresponding transcripts while preserving patient privacy. Acoustic embeddings are extracted directly from speech signals, while textual embeddings are generated from automatically transcribed speech. These modality-specific embeddings are then concatenated to create a combined feature vector and used for downstream classification, without requiring access to raw or sensitive patient data. The proposed approach is evaluated on the ADReSS20 and ADReSSo21 benchmark datasets. Experimental results show that the proposed multimodal framework achieves an CI classification accuracy of 92.4% and consistently outperforms single-modality baselines. Our work establishes a new state-of-the-art for CI identification, with the proposed method demonstrating superior cross-dataset generalization. This advance highlights the power of an LLM-based multimodal framework that fuses linguistic and acoustic data to enable robust, scalable, and non-invasive screening.

认知障碍语音分析多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。