用语音检测阿尔茨海默病,准确率达87%
Swin-BERT: A Feature Fusion System designed for Speech-based Alzheimer's Dementia Detection
- 融合声学与语言特征,引入时序注意力与字符级输入
- 在两个国际数据集上分别达到85.58%和87.32%的F1分数
- 有效分离年龄性别干扰,适合临床辅助诊断场景
语音常被用于构建阿尔茨海默病(AD)自动检测系统,因其声学与语言能力在早期即出现退化。然而,语音中不仅包含与认知状态相关的局部和全局信息,还混杂着年龄、性别等无关因素。本文提出一种名为Swin-BERT的语音基痴呆检测系统:声学部分采用图像中提出的移位窗口多头注意力机制提取局部与全局特征,并将年龄、性别作为额外输入以解耦其影响;语言部分在转录时去除节奏相关特征,改用字符级文本作为词级BERT模型的补充输入以补偿信息损失。最终,Swin-BERT融合声学与语言特征。实验基于国际痴呆检测挑战赛提供的ADReSS与ADReSSo数据集,结果表明,所提声学与语言系统均优于或相当现有方法,而整体系统在两数据集上的F-score分别达到85.58%与87.32%,显著提升检测性能。
原文摘要 · Abstract (English)
Speech is usually used for constructing an automatic Alzheimer's dementia (AD) detection system, as the acoustic and linguistic abilities show a decline in people living with AD at the early stages. However, speech includes not only AD-related local and global information but also other information unrelated to cognitive status, such as age and gender. In this paper, we propose a speech-based system named Swin-BERT for automatic dementia detection. For the acoustic part, the shifted windows multi-head attention that proposed to extract local and global information from images, is used for designing our acoustic-based system. To decouple the effect of age and gender on acoustic feature extraction, they are used as an extra input of the designed acoustic system. For the linguistic part, the rhythm-related information, which varies significantly between people living with and without AD, is removed while transcribing the audio recordings into transcripts. To compensate for the removed rhythm-related information, the character-level transcripts are proposed to be used as the extra input of a word-level BERT-style system. Finally, the Swin-BERT combines the acoustic features learned from our proposed acoustic-based system with our linguistic-based system. The experiments are based on the two datasets provided by the international dementia detection challenges: the ADReSS and ADReSSo. The results show that both the proposed acoustic and linguistic systems can be better or comparable with previous research on the two datasets. Superior results are achieved by the proposed Swin-BERT system on the ADReSS and ADReSSo datasets, which are 85.58\% F-score and 87.32\% F-score respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。