arXiv:2506.01890cs.LGcs.AI2025-06被引 17

通过词级对齐与门控注意力,提升阿尔茨海默病语音文本联合检测精度

CogniAlign: Word-Level Multimodal Speech Alignment with Gated Cross-Attention for Alzheimer's Detection

  • 词级时间对齐同步音频与文本,实现细粒度跨模态融合
  • 在ADReSSo数据集上达到90.36%准确率,优于现有方法
  • 引入停顿标记建模语调特征,适合临床认知障碍筛查研究

阿尔茨海默病等认知障碍的早期检测对及时临床干预和改善患者预后至关重要。本文提出CogniAlign,一种融合语音与文本模态的多模态架构,利用两种非侵入性信息源互补揭示认知健康状态。不同于以往粗粒度模态融合方法,CogniAlign采用词级时间对齐策略,基于转录时间戳同步音频嵌入与对应文本标记,支持细粒度令牌级融合。为此,我们设计了门控交叉注意力融合机制,使音频特征在文本模态优势引导下关注文本表示。同时,通过插入停顿标记并生成静默区间音频嵌入,融入语调线索(如词间停顿),进一步丰富双流信息。在ADReSSo数据集上,模型在留一被试者外部验证中达到87.35%准确率,在五折交叉验证中达90.36%,超越现有最优方法。消融实验验证了对齐策略、注意力融合及语调建模的优势。此外,通过语料分析评估语调特征影响,并使用积分梯度识别模型预测认知状态的关键输入片段。

原文摘要 · Abstract (English)

Early detection of cognitive disorders such as Alzheimer's disease is critical for enabling timely clinical intervention and improving patient outcomes. In this work, we introduce CogniAlign, a multimodal architecture for Alzheimer's detection that integrates audio and textual modalities, two non-intrusive sources of information that offer complementary insights into cognitive health. Unlike prior approaches that fuse modalities at a coarse level, CogniAlign leverages a word-level temporal alignment strategy that synchronizes audio embeddings with corresponding textual tokens based on transcription timestamps. This alignment supports the development of token-level fusion techniques, enabling more precise cross-modal interactions. To fully exploit this alignment, we propose a Gated Cross-Attention Fusion mechanism, where audio features attend over textual representations, guided by the superior unimodal performance of the text modality. In addition, we incorporate prosodic cues, specifically interword pauses, by inserting pause tokens into the text and generating audio embeddings for silent intervals, further enriching both streams. We evaluate CogniAlign on the ADReSSo dataset, where it achieves an accuracy of 87.35% over a Leave-One-Subject-Out setup and of 90.36% over a 5 fold Cross-Validation, outperforming existing state-of-the-art methods. A detailed ablation study confirms the advantages of our alignment strategy, attention-based fusion, and prosodic modeling. Finally, we perform a corpus analysis to assess the impact of the proposed prosodic features and apply Integrated Gradients to identify the most influential input segments used by the model in predicting cognitive health outcomes.

阿尔茨海默病多模态融合语音分析注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。