arXiv:2501.16201eess.AScs.CL2025-01

用语音直接分析识别轻度认知障碍,提升检测效果。

Enhancing and Exploring Mild Cognitive Impairment Detection with W2V-BERT-2.0

  • 直接使用语音特征,避免依赖字幕和时间信息。
  • 新推理逻辑使检测准确率显著优于基线模型。
  • 揭示了说话人偏倚与数据划分对结果的影响。

本研究探索利用多语言音频自监督学习模型,基于TAUKADIAL跨语言数据集检测轻度认知障碍(MCI)。尽管基于语音转录的BERT模型有效,但受限于缺乏转录文本及时间信息。为此,本研究采用W2V-BERT-2.0从语音语句中直接提取特征。提出一种可视化方法以识别用于MCI分类的关键模型层,并设计符合MCI特性的特定推理逻辑。实验显示性能具有竞争力,所提推理逻辑显著提升基线表现。详细分析还揭示了特征中的说话人偏倚问题,以及MCI分类准确率对数据划分的敏感性,为未来研究提供重要洞见。

原文摘要 · Abstract (English)

This study explores a multi-lingual audio self-supervised learning model for detecting mild cognitive impairment (MCI) using the TAUKADIAL cross-lingual dataset. While speech transcription-based detection with BERT models is effective, limitations exist due to a lack of transcriptions and temporal information. To address these issues, the study utilizes features directly from speech utterances with W2V-BERT-2.0. We propose a visualization method to detect essential layers of the model for MCI classification and design a specific inference logic considering the characteristics of MCI. The experiment shows competitive results, and the proposed inference logic significantly contributes to the improvements from the baseline. We also conduct detailed analysis which reveals the challenges related to speaker bias in the features and the sensitivity of MCI classification accuracy to the data split, providing valuable insights for future research.

认知障碍语音分析自监督学习多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。