ASR错误反能提升阿尔茨海默病检测,关键在于挖掘其中隐藏线索。
Beyond Manual Transcripts: The Potential of Automated Speech Recognition Errors in Improving Alzheimer's Disease Detection
- 利用不同ASR模型生成的语音转录文本进行检测研究
- 部分含错误的ASR转录文本准确率高于人工转录
- 提出交叉注意力模型,可解释错误中的潜在诊断信号
近年来,自动语音识别(ASR)技术的进步使得完全自动化阿尔茨海默病(AD)检测成为可能。然而,ASR错误对AD检测的影响仍不明确。本文在ADReSS数据集上对多种ASR模型及其合成语音的转录文本进行了全面研究。实验结果表明,某些含错误的ASR转录文本(ASR-合成语音)在检测准确率上优于人工转录文本(人工-合成语音),提示ASR错误可能包含有助于提升AD检测的有价值线索。此外,我们提出一种基于交叉注意力的可解释性模型,不仅能识别这些线索,且性能优于或相当于基线模型。进一步地,该模型揭示了预训练嵌入中与AD相关的模式。本研究为ASR模型在AD检测中的潜在应用提供了新视角。
原文摘要 · Abstract (English)
Recent breakthroughs in Automatic Speech Recognition (ASR) have enabled fully automated Alzheimer's Disease (AD) detection using ASR transcripts. Nonetheless, the impact of ASR errors on AD detection remains poorly understood. This paper fills the gap. We conduct a comprehensive study on AD detection using transcripts from various ASR models and their synthesized speech on the ADReSS dataset. Experimental results reveal that certain ASR transcripts (ASR-synthesized speech) outperform manual transcripts (manual-synthesized speech) in detection accuracy, suggesting that ASR errors may provide valuable cues for improving AD detection. Additionally, we propose a cross-attention-based interpretability model that not only identifies these cues but also achieves superior or comparable performance to the baseline. Furthermore, we utilize this model to unveil AD-related patterns within pre-trained embeddings. Our study offers novel insights into the potential of ASR models for AD detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。