不依赖语音识别,直接用声谱图检测痴呆早期征兆。
ASR-Agnostic Multimodal Spectrotemporal Modeling for Early Dementia Detection

- 直接处理声谱图,提取频谱能量动态变化作为生物标志物。
- 斯洛伐克数据集达83.9%准确率,英语数据集仅53.2%。
- 多模态融合效果因语种而异,部分情况下反而降低性能。
语音激活与日常活动能力相关的执行、注意和工作记忆过程,可作为认知评估的无创代理指标。然而,多数基于语音的痴呆检测系统依赖语音识别,丢弃录音内部的时间结构,且仅在单一英语语料库上验证,存在已知录音偏差。本文提出一种无需语音识别(ASR-agnostic)的框架,直接在梅尔频谱图上操作。核心贡献是通过连续频谱帧提取时频位移场,捕捉频谱能量迁移模式,作为认知衰退的数字生物标志物。这些特征与CNN-ConvGRU声学嵌入通过学习型交叉注意力机制融合,并由带可学习查询池化的Transformer编码器聚合。复合时间损失强制片段间平滑性与对比一致性。我们在英语DementiaBank、斯洛伐克EWA-DB和西班牙Ivanova语料库上分别训练独立模型,使用针对日常活动能力相关认知域的临床采集协议。斯洛伐克模型达到83.9%准确率,西班牙模型表现优异,而英语基线仅为53.2%,确认了已知数据偏差。跨语言消融研究显示不同融合策略:移除交叉注意力使西班牙性能降至53.7%,低于单模态模型;而斯洛伐克音频编码器单独使用已达93.7%准确率,高于完整模型的83.9%;所有英语配置均接近随机水平。因此,多模态融合的价值取决于语料库特性:当信号分布于多模态时有效,单模态主导时反成负担,无信号时则无效。辅助时间损失收敛至语言无关值,表明架构具备跨语言稳定性。
原文摘要 · Abstract (English)
Speech recruits the same executive, attentional, and working memory processes underlying instrumental activities of daily living, or IADLs, providing a non-invasive proxy for cognitive assessment. Yet most speech-based dementia detection systems depend on transcription, discard within-recording temporal structure, and are validated on a single English corpus with known recording artifacts. We propose an ASR-agnostic framework operating directly on Mel spectrograms. Our key contribution is extracting spectrotemporal displacement fields from consecutive spectrogram frames, capturing shifting spectral energy patterns as digital biomarkers of cognitive decline. These features are fused with CNN-ConvGRU acoustic embeddings via a learned cross-attention mechanism and aggregated using a Transformer encoder with learnable query pooling. A composite temporal loss enforces smoothness and contrastive coherence across segments. We train independent models on English DementiaBank, Slovak EWA-DB, and Spanish Ivanova corpora, using clinical elicitation protocols taxing IADL-relevant cognitive domains. The Slovak model achieves 83.9% accuracy, and Spanish achieves, while the English baseline yields 53.2%, confirming known artifacts. Cross-lingual ablation studies reveal distinct fusion regimes: removing cross-attention collapses Spanish performance to 53.7%, below unimodal models, while the Slovak audio encoder alone outperforms the full model, 93.7% vs. 83.9%, and all English configurations remain near chance. Thus, multimodal fusion's value is corpus-dependent: essential when signal is distributed across modalities, counterproductive when one dominates, and irrelevant when no signal exists. Auxiliary temporal losses converge to language-invariant values, indicating cross-lingual architectural stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。