arXiv:2602.11028cs.CLcs.AI2026-02

用语言特征识别痴呆早期迹象,模型可解释且结果可靠。

Linguistic Indicators of Early Cognitive Decline in the DementiaBank Pitt Corpus: A Statistical and Machine Learning Study

  • 用词性、句法等抽象语言特征建模,不依赖具体词汇
  • 在受试者级验证中仍保持稳定区分能力,准确率超90%
  • 适合临床筛查工具开发,支持透明可解释的决策

背景:自发语言产出中的细微变化是认知衰退的早期指标。识别可语言解释的痴呆标记,有助于构建透明且临床可行的筛查方法。方法:分析DementiaBank Pitt语料库中的自发对话转录文本,采用三种语言表征:原始清洗文本、结合词汇与语法信息的词性(POS)增强表示,以及仅含词性的句法表示。使用逻辑回归与随机森林模型,在两种协议下评估:逐转录本训练测试划分和受试者级五折交叉验证,以避免说话人重叠。通过全局特征重要性分析模型可解释性,并使用曼-惠特尼U检验结合克利夫效应量进行统计验证。结果:在各类表示中,模型表现稳定,即使无词汇内容,句法与语法特征仍具强区分力。受试者级评估结果更为保守但一致,尤其在POS增强与仅POS表示中表现突出。统计分析显示功能词使用、词汇多样性、句子结构与话语连贯性在组间存在显著差异,与机器学习特征重要性结果高度吻合。结论:抽象语言特征能在临床真实条件下捕捉早期认知衰退的稳健标志。结合可解释机器学习与非参数统计验证,本研究支持基于语言学基础特征的透明、可靠的语言认知筛查应用。

原文摘要 · Abstract (English)

Background: Subtle changes in spontaneous language production are among the earliest indicators of cognitive decline. Identifying linguistically interpretable markers of dementia can support transparent and clinically grounded screening approaches. Methods: This study analyzes spontaneous speech transcripts from the DementiaBank Pitt Corpus using three linguistic representations: raw cleaned text, a part-of-speech (POS)-enhanced representation combining lexical and grammatical information, and a POS-only syntactic representation. Logistic regression and random forest models were evaluated under two protocols: transcript-level train-test splits and subject-level five-fold cross-validation to prevent speaker overlap. Model interpretability was examined using global feature importance, and statistical validation was conducted using Mann-Whitney U tests with Cliff's delta effect sizes. Results: Across representations, models achieved stable performance, with syntactic and grammatical features retaining strong discriminative power even in the absence of lexical content. Subject-level evaluation yielded more conservative but consistent results, particularly for POS-enhanced and POS-only representations. Statistical analysis revealed significant group differences in functional word usage, lexical diversity, sentence structure, and discourse coherence, aligning closely with machine learning feature importance findings. Conclusion: The results demonstrate that abstract linguistic features capture robust markers of early cognitive decline under clinically realistic evaluation. By combining interpretable machine learning with non-parametric statistical validation, this study supports the use of linguistically grounded features for transparent and reliable language-based cognitive screening.

语言分析认知衰退机器学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。