通过调控特定神经元,让大模型出现阿尔茨海默症语言症状。
Activation-Guided Neuron Intervention to Induce Alzheimer's-Related Computational Language Phenotypes in a Large Language Model

- 基于临床语言差异定位高激活神经元,调节其输出权重。
- 增强相关神经元导致叙事记忆、流畅性等多维认知功能下降。
- 首次实证神经元可诱发类阿尔茨海默症语言表型,适合认知研究者。
阿尔茨海默病(AD)患者自发语言变化是早期认知障碍的信号,大语言模型(LLMs)可检测此类变化。但仅检测不足以证明模型表征是否功能性参与行为。本文提出一种激活引导干预框架,使用Qwen3-8B模型。框架识别在AD语料中激活率更高的前馈神经元,并通过缩放对应下投影权重来调节生成过程中的贡献。由此产生九种不同干预方向、强度和范围的修改版本。原始模型与修改模型均完成12轮神经心理量表测试,由盲评人员及计算语言学指标评估。增强AD相关神经元导致故事回忆、词汇流畅性、工作记忆、程序化话语、场景构建和指代消解等能力呈梯度下降;抑制则基本维持性能,并在部分指标上有所提升。增强还降低词汇意外性、思想密度、句法复杂度和话语量,与人类AD语言特征广泛一致。结果表明,仅基于临床语言差异识别的神经元即可影响多个认知领域行为,为阿尔茨海默症相关计算语言表型提供概念验证,并建立可控实验框架以探究语言与更广泛认知失调的关联。
原文摘要 · Abstract (English)
Changes in spontaneous speech provide an early signal of cognitive dysfunction in Alzheimer's disease (AD) that large language models (LLMs) can detect. However, detection alone cannot establish whether the underlying model representations contribute functionally to behavior. We introduce an activation-guided intervention framework using Qwen3-8B. The framework identifies feed-forward neurons with higher activation rates for AD than control transcripts and modulates their output contributions during generation by scaling the corresponding down-projection weights. This yielded nine edited variants differing in intervention direction, magnitude, and scope. The original and edited models completed the same 12-turn neuropsychological battery, assessed through blinded human ratings and computational linguistic measures. Amplifying AD-associated neurons produced graded impairments in story recall, verbal fluency, working memory, procedural discourse, scene construction, and coreference resolution. Attenuation largely preserved performance and selectively improved several outcomes. Amplification also reduced lexical surprisal, idea density, syntactic complexity, and discourse quantity, broadly paralleling changes reported in human AD speech. These findings show that neurons identified solely from clinical language differences can influence behavior across multiple cognitive domains, providing proof of concept for an AD-related computational phenotype and a controlled framework for experimentally examining links between language and broader cognitive dysfunction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。