arXiv:2608.06429cs.CLcs.LG2026-08

通过反向还原语言模型的损伤参数,揭示了变压器层间的功能冗余。

Recovering Lesion Parameters from Aphasic Picture Naming Error Profiles in Large Language Models

  • 用多任务神经网络从命名错误模式反推模型损伤参数
  • 81.4%的反事实验证成功率,表明恢复参数可复现原行为
  • 发现低层损伤难以精确还原,反映层间功能冗余

大型语言模型(LLMs)的可解释性方法描述内部状态,但未直接检验该状态是否足以导致观察到的行为。此前研究中,我们对LLaVA-Vicuna 13B进行损伤,生成图片命名错误模式,其表现与中风患者个体症状相似。本文提出逆问题:给定错误模式,能否恢复出对应的损伤参数?共测试4,840种配置,损伤参数包括层索引、修改比例和噪声方差;错误模式采用七类临床分类法(正确、语义错误、无关、形式错误、混合、新词、无应答)。训练多任务神经网络实现从错误模式到扰动参数的映射。结果部分成功:在10个独立训练的逆模型中,修改比例与噪声方差可被恢复,层索引仅在邻域内可还原。反事实验证中,使用恢复参数扰动新模型实例,在81.4%情况下重现目标行为。这一低层不可精确定位但高反事实保真度的现象,暗示变压器层间存在功能冗余,而传统可解释性方法无法捕捉此特性。作为分布外测试,将模型应用于278名中风患者的命名错误数据,恢复参数具有综合征判别能力,尤其对扰动强度敏感,表明模型泛化能力超越训练分布。反事实验证为大模型可解释性声明提供通用框架。

原文摘要 · Abstract (English)

Interpretability methods for large language models (LLMs) describe internal state but do not directly test whether that state is causally sufficient to produce the observed behavior. In earlier work, we lesioned LLMs to produce error profiles in picture naming, a central task for assessing aphasia, and found that specific lesions produced errors resembling those of individual stroke survivors. Here we ask the inverse question: given an error profile, can the lesion parameters that produced it be recovered, and what does this inverse problem reveal about transformer computation? Lesions in LLaVA-Vicuna 13B were parameterized by layer index, modification percentage, and noise sigma across 4,840 configurations, and error profiles were characterized by a seven-category clinical taxonomy (correct, semantic, unrelated, formal, mixed, neologism, no-response). We trained a multi-task neural network to map error profiles back to perturbation parameters. The problem admitted a partial solution: across 10 independently trained inverse models, modification percentage and noise sigma were recoverable, whereas layer index was recoverable only within a neighborhood. In counterfactual validation, a fresh model instance perturbed with the recovered parameters reproduced the target behavior in 81.4% of cases. This dissociation between low layer recovery and high counterfactual fidelity is consistent with functional redundancy across transformer layers, a property not captured by standard interpretability methods. As an out-of-distribution test, we applied the trained model to picture-naming error profiles from 278 stroke survivors; recovered parameters were syndrome-discriminative, most strongly for perturbation intensity, indicating generalization beyond the training distribution. Counterfactual validation provides a general framework for LLM interpretability claims beyond inverse mapping.

可解释性语言模型神经机制反事实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。