通过语音与文本联合建模,提升阿尔茨海默病早期检测准确率
A Multimodal Framework for Dementia Detection via Linguistic and Acoustic Representation Learning

- 端到端融合语音与转录文本特征,增强模态间关联性
- 在ADReSS与PROCESS-2数据集上达到91.7%与86.4%准确率
- 适合医疗AI、认知障碍筛查领域的研究者与开发者
阿尔茨海默病(AD)是一种进行性神经退行性疾病,是痴呆的主要原因,影响记忆、推理、沟通及日常功能。早期诊断至关重要,及时干预可能延缓认知衰退并改善患者护理。近期研究表明,自发性言语中包含与痴呆相关的语言和声学生物标志物。然而,现有方法通常依赖独立训练的模态特定模型、特征拼接、集成方法或基于注意力的融合机制,未能显式最大化语音与转录文本表示之间的依赖关系。本文提出一种多模态深度学习框架,用于自动痴呆检测,以端到端方式联合利用语音与转录信息。具体而言,语音记录被分割为10秒片段,通过预训练HuBERT模型提取上下文感知的声学表示;为更好捕捉有信息量的时间特征,采用注意力统计池化聚合帧级声学嵌入。对于文本模态,使用预训练BERT模型编码转录内容,取[CLS] token表示作为语言嵌入。声学与语言表示通过基于注意力的音频-文本融合(AT-Fusion)机制结合。此外,引入MINE目标以最大化模态间互信息,提升多模态表示对齐。融合后的多模态表示最终用于痴呆分类。在公开的ADReSS Challenge与PROCESS-2数据集上的实验表明,所提方法在语音基痴呆评估中具有有效性与鲁棒性。
原文摘要 · Abstract (English)
Alzheimer's disease (AD) is a progressive neurodegenerative disorder and the leading cause of dementia, affecting memory, reasoning, communication, and daily functioning. Early diagnosis is particularly important, as timely intervention may help slow cognitive decline and improve patient care. Recent studies have demonstrated that spontaneous speech contains valuable linguistic and acoustic biomarkers associated with dementia. However, existing approaches often rely on independently trained modality-specific models, feature concatenation strategies, ensemble methods, or attention-based fusion mechanisms that do not explicitly maximize the dependency between speech and transcript representations. In this work, we propose a multimodal deep learning framework for automatic dementia detection that jointly exploits speech and transcript information in an end-to-end trainable manner. Specifically, speech recordings are divided into 10-second segments and passed through a pre-trained HuBERT model to extract contextualized acoustic representations. To better capture informative temporal speech characteristics, attentive statistics pooling is employed to aggregate frame-level acoustic embeddings. For the textual modality, transcripts are encoded using a pre-trained BERT model, where the [CLS] token representation is used as the linguistic embedding. The acoustic and textual representations are subsequently combined using an attention-based Audio-Text Fusion (AT-Fusion) mechanism. In addition, we introduce a MINE objective to maximize the mutual information between modalities and improve multimodal representation alignment. The fused multimodal representation is finally used for dementia classification. Experiments conducted on the publicly available ADReSS Challenge and PROCESS-2 dataset demonstrate the effectiveness and robustness of the proposed approach for speech-based dementia assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。