用动态音频时序建模提升阿尔茨海默病语音检测准确率
Temporal-Aware Iterative Speech Model for Dementia Detection
- 将语音谱图视为连续帧,用卷积GRU捕捉声学特征逐帧演化
- 通过交叉注意力对齐音调与停顿,揭示认知衰退的语音缺陷
- 直接处理原始音频,不依赖语音识别,在DementiaBank上达AUC 0.839
深度学习在处理长序列时常因计算复杂度受限。当前基于语音的阿尔茨海默病自动检测多依赖静态、时间无关特征或聚合语言内容,难以捕捉语音生成中细微渐进的退化过程,忽略关键的动态时序模式。本文提出TAI-Speech,一种时间感知迭代框架,动态建模自发语音以实现痴呆检测。核心创新包括:1)受光流启发的迭代细化,将频谱图视为序列帧,利用卷积GRU捕捉声学特征的细粒度帧间演变;2)基于交叉注意力的韵律对齐,动态关联谱特征与音调、停顿等韵律模式,构建与日常生活能力下降(IADL)相关的语音生产缺陷表示。TAI-Speech自适应建模每个语句的时序演化,增强认知标志物检测。在DementiaBank数据集上的实验表明,其达到0.839的AUC和80.6%的准确率,优于无需ASR的文本基线方法。本工作为自动化认知评估提供更灵活鲁棒的解决方案,直接作用于原始音频的动态特性。
原文摘要 · Abstract (English)
Deep learning systems often struggle with processing long sequences, where computational complexity can become a bottleneck. Current methods for automated dementia detection using speech frequently rely on static, time-agnostic features or aggregated linguistic content, lacking the flexibility to model the subtle, progressive deterioration inherent in speech production. These approaches often miss the dynamic temporal patterns that are critical early indicators of cognitive decline. In this paper, we introduce TAI-Speech, a Temporal Aware Iterative framework that dynamically models spontaneous speech for dementia detection. The flexibility of our method is demonstrated through two key innovations: 1) Optical Flow-inspired Iterative Refinement: By treating spectrograms as sequential frames, this component uses a convolutional GRU to capture the fine-grained, frame-to-frame evolution of acoustic features. 2) Cross-Attention Based Prosodic Alignment: This component dynamically aligns spectral features with prosodic patterns, such as pitch and pauses, to create a richer representation of speech production deficits linked to functional decline (IADL). TAI-Speech adaptively models the temporal evolution of each utterance, enhancing the detection of cognitive markers. Experimental results on the DementiaBank dataset show that TAI-Speech achieves a strong AUC of 0.839 and 80.6\% accuracy, outperforming text-based baselines without relying on ASR. Our work provides a more flexible and robust solution for automated cognitive assessment, operating directly on the dynamics of raw audio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。