arXiv:2501.18731cs.LGcs.CL2025-01被引 16

用语音特征自动筛查认知障碍,准确率超69%。

Evaluating Spoken Language as a Biomarker for Automated Screening of Cognitive Impairment

  • 基于语言特征的随机森林模型识别痴呆症,灵敏度69.4%。
  • 在真实场景数据中仍保持70%灵敏度,特异性提升13%。
  • 适合居家认知健康监测与高风险人群早期筛查。

及时准确评估认知障碍仍是重大未满足需求。语音生物标志物为自动化筛查提供了可扩展、无创、低成本的解决方案。然而,机器学习在临床应用中的效用受限于可解释性及在真实语音数据集上的泛化能力。本研究评估了可解释机器学习在阿尔茨海默病及相关痴呆症(ADRD)筛查与严重程度预测中的表现,使用基准DementiaBank语音数据集(N = 291,女性占比64%,平均年龄69.8岁,标准差8.6)。在院内采集的试点数据上验证泛化能力(N = 22,女性59%,平均年龄76.2岁,标准差8.0)。通过风险分层实现可操作的分诊,并分析语言特征重要性。结果显示,基于语言特征训练的随机森林模型在ADRD检测中平均灵敏度达69.4%(95%置信区间66.4–72.5),特异性为83.3%(78.0–88.7)。在试点数据上,灵敏度为70.0%(58.0–82.0),特异性为52.5%(39.3–65.7)。对于迷你精神状态检查(MMSE)得分预测,随机森林回归器平均绝对误差为3.7(3.7–3.8),试点数据上为3.3(3.1–3.5)。风险分层使测试集特异性提高13%,为临床分诊提供路径。与ADRD相关的语言特征包括:代词和副词使用增加、言语不流畅性上升、分析性思维减少、词汇多样性降低、反映心理完成状态的词汇减少。该模型显示有潜力集成至家庭对话技术中,持续监测认知健康并分诊高风险个体,实现早期筛查与干预。

原文摘要 · Abstract (English)

Timely and accurate assessment of cognitive impairment remains a major unmet need. Speech biomarkers offer a scalable, non-invasive, cost-effective solution for automated screening. However, the clinical utility of machine learning (ML) remains limited by interpretability and generalisability to real-world speech datasets. We evaluate explainable ML for screening of Alzheimer's disease and related dementias (ADRD) and severity prediction using benchmark DementiaBank speech (N = 291, 64% female, 69.8 (SD = 8.6) years). We validate generalisability on pilot data collected in-residence (N = 22, 59% female, 76.2 (SD = 8.0) years). To enhance clinical utility, we stratify risk for actionable triage and assess linguistic feature importance. We show that a Random Forest trained on linguistic features for ADRD detection achieves a mean sensitivity of 69.4% (95% confidence interval (CI) = 66.4-72.5) and specificity of 83.3% (78.0-88.7). On pilot data, this model yields a mean sensitivity of 70.0% (58.0-82.0) and specificity of 52.5% (39.3-65.7). For prediction of Mini-Mental State Examination (MMSE) scores, a Random Forest Regressor achieves a mean absolute MMSE error of 3.7 (3.7-3.8), with comparable performance of 3.3 (3.1-3.5) on pilot data. Risk stratification improves specificity by 13% on the test set, offering a pathway for clinical triage. Linguistic features associated with ADRD include increased use of pronouns and adverbs, greater disfluency, reduced analytical thinking, lower lexical diversity, and fewer words that reflect a psychological state of completion. Our predictive modelling shows promise for integration with conversational technology at home to monitor cognitive health and triage higher-risk individuals, enabling early screening and intervention.

认知筛查语音生物标志物机器学习居家监测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。