arXiv:2603.18239q-bio.QMcs.CL2026-03

高精度语音识别能显著提升阿尔茨海默病筛查效果,简单模型即可达到良好性能。

Impact of automatic speech recognition quality on Alzheimer's disease detection from spontaneous speech: a reproducible benchmark study with lexical modeling and statistical validation

  • 用Whisper ASR提取词汇特征,结合TF-IDF和线性SVM进行分类。
  • Whisper-small比Whisper-base表现更好,平衡准确率超0.7850。
  • 语音识别质量比模型复杂度更重要,适合临床AI系统设计参考。

从自发性语言中早期检测阿尔茨海默病已成为一种有前景的无创筛查方法。然而,自动语音识别(ASR)质量对下游语言建模的影响仍不明确。本研究在ADReSSo 2021诊断数据集上,利用Whisper ASR转录文本,基于词汇特征使用可解释机器学习模型(逻辑回归、线性支持向量机)与TF-IDF表示,在重复5×5分层交叉验证下评估。结果表明,转录质量对分类性能有统计学显著影响:使用Whisper-small转录的模型始终优于Whisper-base,线性SVM平衡准确率超过0.7850。配对统计检验确认差异显著。值得注意的是,分类器复杂度对性能变化的影响小于ASR转录质量。特征分析显示,认知正常者语言更精确,描述物体与场景;而阿尔茨海默病患者语言更模糊,频繁使用话语标记与停顿。这些发现表明,高质量的ASR可使简单的可解释词汇模型在无需显式声学建模的情况下实现竞争性检测性能。本研究提供可复现的基准流程,强调了在临床语音人工智能系统中选择ASR的重要性。

原文摘要 · Abstract (English)

Early detection of Alzheimer's disease from spontaneous speech has emerged as a promising non-invasive screening approach. However, the influence of automatic speech recognition (ASR) quality on downstream clinical language modeling remains insufficiently understood. In this study, we investigate Alzheimer's disease detection using lexical features derived from Whisper ASR transcripts on the ADReSSo 2021 diagnosis dataset. We evaluate interpretable machine-learning models, including Logistic Regression and Linear Support Vector Machines, using TF-IDF text representations under repeated 5x5 stratified cross-validation. Our results demonstrate that transcript quality has a statistically significant impact on classification performance. Models trained on Whisper-small transcripts consistently outperform those using Whisper-base transcripts, achieving balanced accuracy above 0.7850 with Linear SVM. Paired statistical testing confirms that the observed improvements are significant. Importantly, classifier complexity contributes less to performance variation than ASR transcription quality. Feature analysis reveals that cognitively normal speakers produce more semantically precise object- and scene-descriptive language, whereas Alzheimer's speech is characterized by vagueness, discourse markers, and increased hesitation patterns. These findings suggest that high-quality ASR can enable simple, interpretable lexical models to achieve competitive Alzheimer's detection performance without explicit acoustic modeling. The study provides a reproducible benchmark pipeline and highlights ASR selection as a critical modeling decision in clinical speech-based artificial intelligence systems.

阿尔茨海默病语音识别临床AI词汇建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。