通过分层融合声学与语言特征,提升认知状态分类准确率
Layer-Aware Early Fusion of Acoustic and Linguistic Embeddings for Cognitive Status Classification
- 按编码器层深度选择性融合语音与文本嵌入,实现多模态协同
- 中层(约第8-10层)融合效果最佳,最高F1达特定组合
- 适合关注老年认知障碍早期诊断的研究者使用
言语包含反映认知衰退的声学与语言模式,仅描述单一模态的模型难以捕捉其复杂性。本研究探讨在不同编码器层深度下,语音与其对应转录文本嵌入的早期融合(EF)对认知状态分类的影响。基于从DementiaBank获取的录音数据集(1,629名被试,包括认知正常对照组CN、轻度认知障碍MCI及阿尔茨海默病及相关痴呆ADRD),我们从wav2vec 2.0或Whisper结合DistilBERT或RoBERTa的不同内部层提取帧对齐嵌入。训练了单模态、早期融合(EF)与晚期融合(LF)模型,采用Transformer分类器进行优化,并在10个随机种子下评估性能。结果表明,中层(约第8–10层)表现最优,其中Whisper + RoBERTa第9层取得最高F1,Whisper + DistilBERT第10层获得最低对数损失。纯声学模型始终优于纯文本模型。早期融合增强声学特征判别力,而晚期融合改善概率校准。编码器层的选择显著影响临床多模态协同效果。
原文摘要 · Abstract (English)
Speech contains both acoustic and linguistic patterns that reflect cognitive decline, and therefore models describing only one domain cannot fully capture such complexity. This study investigates how early fusion (EF) of speech and its corresponding transcription text embeddings, with attention to encoder layer depth, can improve cognitive status classification. Using a DementiaBank-derived collection of recordings (1,629 speakers; cognitively normal controls$\unicode{x2013}$CN, Mild Cognitive Impairment$\unicode{x2013}$MCI, and Alzheimer's Disease and Related Dementias$\unicode{x2013}$ADRD), we extracted frame-aligned embeddings from different internal layers of wav2vec 2.0 or Whisper combined with DistilBERT or RoBERTa. Unimodal, EF and late fusion (LF) models were trained with a transformer classifier, optimized, and then evaluated across 10 seeds. Performance consistently peaked in mid encoder layers ($\sim$8$\unicode{x2013}$10), with the single best F1 at Whisper + RoBERTa layer 9 and the best log loss at Whisper + DistilBERT layer 10. Acoustic-only models consistently outperformed text-only variants. EF boosts discrimination for genuinely acoustic embeddings, whereas LF improves probability calibration. Layer choice critically shapes clinical multimodal synergy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。