用双阶段网络分析语音,提升阿尔茨海默病早期检测准确率
A Dual-Stage Time-Context Network for Speech-Based Alzheimer's Disease Detection
- 分段处理长语音,先提取局部声学特征再融合全局对话上下文
- 在ADReSSo数据集上达83.10%准确率和83.15%F1值,优于现有方法
- 适合关注语音诊断与神经退行性疾病早期筛查的研究者
阿尔茨海默病(AD)是一种进行性神经退行性疾病,导致记忆与沟通能力不可逆下降。通过语音分析实现早期检测对延缓疾病进展至关重要。然而,现有方法多依赖预训练声学模型提取特征,难以同时建模长时语音中的局部与全局模式。本文提出双阶段时序-上下文网络(DSTC-Net),将局部声学特征与全局对话上下文融合于长时录音中。首先将长语音划分为固定长度片段以降低计算开销并保留局部时间细节;随后,通过帧级注意力的双向长短期记忆网络(BiLSTM)提取增强的局部特征;最后,利用基于卷积的跨片段上下文注意力(CSCA)模块实现全片段的自适应全局模式统一。在ADReSSo数据集上的大量实验表明,DSTC-Net超越当前最优模型,达到83.10%准确率和83.15% F1值。
原文摘要 · Abstract (English)
Alzheimer's disease (AD) is a progressive neurodegenerative disorder that leads to irreversible cognitive decline in memory and communication. Early detection of AD through speech analysis is crucial for delaying disease progression. However, existing methods mainly use pre-trained acoustic models for feature extraction but have limited ability to model both local and global patterns in long-duration speech. In this letter, we introduce a Dual-Stage Time-Context Network (DSTC-Net) for speech-based AD detection, integrating local acoustic features with global conversational context in long-duration recordings.We first partition each long-duration recording into fixed-length segments to reduce computational overhead and preserve local temporal details.Next, we feed these segments into an Intra-Segment Temporal Attention (ISTA) module, where a bidirectional Long Short-Term Memory (BiLSTM) network with frame-level attention extracts enhanced local features.Subsequently, a Cross-Segment Context Attention (CSCA) module applies convolution-based context modeling and adaptive attention to unify global patterns across all segments.Extensive experiments on the ADReSSo dataset show that our DSTC-Net outperforms state-of-the-art models, reaching 83.10% accuracy and 83.15% F1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。