用方言识别模型分析语音,端到端判断阿尔茨海默病和轻度认知障碍。
VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge
- 基于语音基础模型,利用方言特征检测认知障碍。
- 在ADReSS2020和ADReSSo2021上分别达到87.5%和85.9%准确率。
- 仅用语音即可超越多模态融合方法,适合临床筛查场景。
本文提出一种新视角:通过整合能显式识别方言的语音基础模型,从语音中进行认知障碍分类。研究发现,阿尔茨海默病(AD)或轻度认知障碍(MCI)患者常表现出类似方言音变的可测量语音特征,如语速变慢、音节延长。基于此,我们构建了端到端框架VoxCog,利用预训练方言模型在无需文本或图像等额外模态的情况下检测AD或MCI。在多个多语言数据集上的实验表明,以方言分类器初始化语音基础模型可持续提升预测性能。训练模型在ADReSS 2020挑战赛和ADReSSo 2021挑战赛测试集上分别取得87.5%和85.9%的准确率,表现优于依赖多模态融合或大模型的方法。
原文摘要 · Abstract (English)
In this work, we present a novel perspective on cognitive impairment classification from speech by integrating speech foundation models that explicitly recognize speech dialects. Our motivation is based on the observation that individuals with Alzheimer's Disease (AD) or mild cognitive impairment (MCI) often produce measurable speech characteristics, such as slower articulation rate and lengthened sounds, in a manner similar to dialectal phonetic variations seen in speech. Building on this idea, we introduce VoxCog, an end-to-end framework that uses pre-trained dialect models to detect AD or MCI without relying on additional modalities such as text or images. Through experiments on multiple multilingual datasets for AD and MCI detection, we demonstrate that model initialization with a dialect classifier on top of speech foundation models consistently improves the predictive performance of AD or MCI. Our trained models yield similar or often better performance compared to previous approaches that ensembled several computational methods using different signal modalities. Particularly, our end-to-end speech-based model achieves 87.5% and 85.9% accuracy on the ADReSS 2020 challenge and ADReSSo 2021 challenge test sets, outperforming existing solutions that use multimodal ensemble-based computation or LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。