不依赖语音转录,仅用原始音频就能筛查阿尔茨海默病。
Transcript-Free Lightweight Detection of Alzheimer's Disease from Spontaneous Speech Using Handcrafted MFCC-Dominant Acoustic Biomarkers
- 用手工提取的99个声学特征,结合轻量SVM模型分析语音
- 在30次独立测试中平均AUC达0.674,最高达0.742
- 适合资源有限场景,为实际部署提供可行基线
阿尔茨海默病(AD)早期检测仍具挑战性,尤其当神经影像成本高或语言依赖工具不可用时。自发语音是一种非侵入性信号,但现有方法多依赖语音转录或计算密集型深度模型。本文提出一种仅基于音频的简化基线方法,使用来自DementiaBank Pitt语料库的176段Cookie Theft录音(88例AD患者,88例对照者)。通过WebRTC语音活动检测(VAD)分离语音与非语音部分。提取99个手工设计的声学-时序特征,包括停顿、流畅性统计量、频谱/韵律描述符以及包含Δ和ΔΔ的MFCC摘要。采用严格的说话人无关的GroupShuffleSplit进行评估,共30次迭代。使用径向基函数核的轻量SVM平均AUC为0.674;例如某次划分的AUC为0.742,准确率为0.657。此外,通过随机森林重要性排序的前20个紧凑特征进行探索性分析,结果可能因未嵌套训练而偏乐观(AUC 0.719),未用于主结论。结果表明,无需转录的频谱-时序与流畅性线索可支持从原始音频实现说话人无关的阿尔茨海默病筛查,为面向应用的研究奠定实用基础。
原文摘要 · Abstract (English)
It is still hard to find Alzheimer's disease (AD) early, especially when neuroimaging is expensive or tools that depend on language are not available. Spontaneous speech provides a non-invasive signal; however, numerous current methodologies depend on transcripts/ASR or computationally intensive deep models. We offer a simple, audio-only baseline for detecting AD using 176 Cookie Theft recordings from the DementiaBank Pitt corpus (88 AD, 88 controls). WebRTC voice activity detection (VAD) is used to separate speech from non-speech. We take out 99 hand-crafted acoustic-temporal features, including pause and fluency statistics, spectral/prosodic descriptors, and MFCC summaries with Δ and ΔΔ. Evaluation is performed using a stringent speaker-independent GroupShuffleSplit,documenting performance across 30 iterations. A lightweight SVM with an RBF kernel gets an average AUC of 0.674 across runs. For example, a single split has an AUC of 0.742 and an accuracy of 0.657. We also present an exploratory compact-feature analysis utilizing a Top-20 subset ranked by Random Forest importance; since selection is not nested within training splits, these results may be overly optimistic and are not employed for primary conclusions (AUC 0.719). The results indicate that transcript-free spectro-temporal and fluency-related cues can facilitate speaker-independent Alzheimer's disease screening from raw audio, establishing a practical foundation for deployment-oriented research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。