arXiv:2505.17477cs.LGcs.SD2025-05被引 2

用神经网络逆向生成阿尔茨海默病语音样本,提升诊断准确率。

Reverse-Speech-Finder: A Neural Network Backtracking Architecture for Generating Alzheimer's Disease Speech Samples and Improving Diagnosis Performance

  • 通过反向追踪最可能预测阿尔茨海默病的神经元,定位关键语音标记。
  • 在真实数据稀缺情况下,生成新语音样本,使准确率提升3.5%、F1得分提高3.2%。
  • 适合关注语言障碍与早期诊断的临床研究者和医疗AI开发者。

本研究提出逆向语音发现器(Reverse-Speech-Finder, RSF),一种创新的神经网络反向追踪架构,旨在通过语音分析提升阿尔茨海默病(AD)诊断性能。利用预训练大语言模型,RSF识别并利用最具可能预测AD的语音标记(即最可能语音标记,MPMs),解决真实AD语音样本稀缺及现有模型可解释性差的问题。其核心创新包括:首先,基于最可能预测AD的神经元(MPNs)激活概率,反推最可能触发这些神经元的语音标记;其次,采用输入层语音标记表示,实现从MPNs回溯至最可能语音标记(MPTs);最后,设计新型反向追踪方法,从输出层回溯至输入层,识别出对应的MPTs与MPMs,揭示了新的AD检测语音特征。实验表明,相较SHAP与Integrated Gradients等传统方法,RSF在准确率上提升3.5%,F1-score提升3.2%。通过生成包含新标记的语音数据,RSF不仅缓解真实数据不足问题,还显著增强诊断模型的鲁棒性与准确性。该成果为基于语音的阿尔茨海默病检测提供了变革性工具,深化对语言缺陷的理解,并推动非侵入性早期干预策略发展。

原文摘要 · Abstract (English)

This study introduces Reverse-Speech-Finder (RSF), a groundbreaking neural network backtracking architecture designed to enhance Alzheimer's Disease (AD) diagnosis through speech analysis. Leveraging the power of pre-trained large language models, RSF identifies and utilizes the most probable AD-specific speech markers, addressing both the scarcity of real AD speech samples and the challenge of limited interpretability in existing models. RSF's unique approach consists of three core innovations: Firstly, it exploits the observation that speech markers most probable of predicting AD, defined as the most probable speech-markers (MPMs), must have the highest probability of activating those neurons (in the neural network) with the highest probability of predicting AD, defined as the most probable neurons (MPNs). Secondly, it utilizes a speech token representation at the input layer, allowing backtracking from MPNs to identify the most probable speech-tokens (MPTs) of AD. Lastly, it develops an innovative backtracking method to track backwards from the MPNs to the input layer, identifying the MPTs and the corresponding MPMs, and ingeniously uncovering novel speech markers for AD detection. Experimental results demonstrate RSF's superiority over traditional methods such as SHAP and Integrated Gradients, achieving a 3.5% improvement in accuracy and a 3.2% boost in F1-score. By generating speech data that encapsulates novel markers, RSF not only mitigates the limitations of real data scarcity but also significantly enhances the robustness and accuracy of AD diagnostic models. These findings underscore RSF's potential as a transformative tool in speech-based AD detection, offering new insights into AD-related linguistic deficits and paving the way for more effective non-invasive early intervention strategies.

阿尔茨海默病语音分析生成模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。