arXiv:2512.13685cs.CL2025-12

通过语义保持的文本变换,验证语言模型可仅凭深层语义识别阿尔茨海默病。

Beyond surface form: A pipeline for semantic analysis in Alzheimer's Disease detection from spontaneous speech

  • 用语法词汇改写保持语义,分离表面形式影响。
  • 模型在变换后仍保持高分类准确率,宏平均F1仅小幅下降。
  • 证明语义信息足以检测阿尔茨海默病,适合早期筛查研究者。

阿尔茨海默病(AD)是一种进行性神经退行性疾病,损害认知能力。语言相关变化可通过语言评估任务(如图片描述)输出自动识别。语言模型有望作为AD筛查工具基础,但其可解释性有限,难以区分真实语言标志与表面文本模式。为解决此问题,我们研究表面形式变化对分类性能的影响,旨在评估语言模型对潜在语义指标的表征能力。提出新方法:通过改变句法和词汇而保留语义内容来转换文本。变换显著改变结构与词项,表现为低BLEU和chrF得分,但语义相似度高,有效隔离语义影响。结果显示,模型在变换文本上的表现与原始文本几乎一致,宏平均F1仅微小偏差。此外,我们检验图片描述是否足够重建原图,发现基于图像的变换引入大量噪声,降低分类精度。该方法提供了一种分析影响模型预测特征的新路径,可消除虚假相关性。结果表明,仅依赖语义信息,语言模型仍能检测AD。本工作揭示了难以察觉的语义损伤可被识别,弥补语言退化中被忽视的方面,为早期检测系统开辟新途径。

原文摘要 · Abstract (English)

Alzheimer's Disease (AD) is a progressive neurodegenerative condition that adversely affects cognitive abilities. Language-related changes can be automatically identified through the analysis of outputs from linguistic assessment tasks, such as picture description. Language models show promise as a basis for screening tools for AD, but their limited interpretability poses a challenge in distinguishing true linguistic markers of cognitive decline from surface-level textual patterns. To address this issue, we examine how surface form variation affects classification performance, with the goal of assessing the ability of language models to represent underlying semantic indicators. We introduce a novel approach where texts surface forms are transformed by altering syntax and vocabulary while preserving semantic content. The transformations significantly modify the structure and lexical content, as indicated by low BLEU and chrF scores, yet retain the underlying semantics, as reflected in high semantic similarity scores, isolating the effect of semantic information, and finding models perform similarly to if they were using the original text, with only small deviations in macro-F1. We also investigate whether language from picture descriptions retains enough detail to reconstruct the original image using generative models. We found that image-based transformations add substantial noise reducing classification accuracy. Our methodology provides a novel way of looking at what features influence model predictions, and allows the removal of possible spurious correlations. We find that just using semantic information, language model based classifiers can still detect AD. This work shows that difficult to detect semantic impairment can be identified, addressing an overlooked feature of linguistic deterioration, and opening new pathways for early detection systems.

阿尔茨海默病语义分析语音识别语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。