arXiv:2606.30675eess.AScs.AI2026-06中稿 · INTERSPEECH 2026

用语音模型同时提取声音和语言特征,提升阿尔茨海默病早期检测准确率。

Listening Between the Lines: Joint Learning of ASR Embeddings and LLM-Augmented Linguistics for Dementia Detection

论文配图:Listening Between the Lines: Joint Learning of ASR Embeddings and LLM-Augmented Linguistics for Dementia Detection
图 1 · 摘自论文原文
  • 用Whisper同时获取语音声学特征和文字转录内容
  • 结合大模型分析词汇多样性等语言特征,F1达90.14%
  • 适合做临床辅助诊断或语音分析研究者参考

通过语音分析进行阿尔茨海默病的早期检测提供了一种无创筛查方式,但同时捕捉声学与语言生物标志物仍具挑战。本文提出一种多模态框架,利用Whisper实现双重提取:从编码器输出获取声学表征,通过自动语音识别(ASR)获得转录文本。声学路径采用带注意力池化的时序网络,将可变长度序列聚合为固定维度嵌入;语言路径则通过大语言模型(LLM)提示,提取涵盖词汇多样性、句法复杂度、语义连贯性及话语模式的可解释特征。一个门控融合网络整合双模态信息。在ADReSS和ADReSSo数据集上,该方法分别取得89.47%和90.14%的F1分数,证明了声学与LLM增强语言特征的有效融合。消融实验显示,多模态融合始终优于单一模态。

原文摘要 · Abstract (English)

Early detection of dementia through speech analysis offers a non-invasive screening alternative, but capturing both acoustic and linguistic biomarkers remains challenging. We propose a multimodal framework leveraging Whisper for dual-purpose extraction: acoustic representations from encoder outputs and transcripts via automatic speech recognition (ASR). For the acoustic pathway, temporal networks with attention pooling aggregate variable-length sequences into fixed-dimensional embeddings. For the linguistic pathway, we prompt a large language model (LLM) to extract interpretable features spanning lexical diversity, syntactic complexity, semantic coherence, and discourse patterns. A gated fusion network integrates both modalities. On ADReSS and ADReSSo, our method achieves F1-scores of 89.47% and 90.14%, demonstrating effective integration of acoustic and LLM-augmented linguistic features. Ablation shows that multimodal fusion consistently outperforms either modality alone.

语音分析阿尔茨海默病多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。