arXiv:2608.03842cs.CLcs.LG2026-08

不同层在抗扰性、因果性和修复能力上表现不一,揭示了模型内部故障机制的复杂性。

Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling

论文配图:Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling
图 1 · 摘自论文原文
  • 通过三类指标区分模型各层对扰动的敏感度、因果贡献与修复潜力
  • 发现早期层的修复适配器会破坏下游计算,导致诊断位置反而最差
  • 建议优先在深层部署修复适配器,并提供无需训练的预筛选方法

当语言模型在存在拼写错误、光学字符识别噪声或同音词的输入上失效时,‘哪一层负责’可从三个角度理解:表征偏离最严重的层(敏感性)、恢复干净激活后预测得以恢复的层(因果性),以及小适配器能有效修复损伤的层(补偿能力)——我们发现这三个层图谱相互分离。在五模型面板中,识别出两种传播模式:脉冲-抑制型(Phi-3.5、Gemma-2-9B)和后期累积型(Llama-3、Mistral、Qwen2.5-7B)。在满足80%身份补丁门限的两个模型上,敏感性与因果性呈强负相关(rho = -0.72 至 -0.88)。Qwen2.5系列(1.5B 到 14B)的家族内缩放显示后期累积特征随规模单调增强,另一家族亦得验证。我们提出级联中断作为分离机制:放置于因果关键早期层的适配器会破坏完整下游计算,使诊断标记点成为最差的适配器位置。在四个模型(3.8-8B)上进行固定框架的层扫描,确认其核心预测在链式思维数学题GSM8K上成立——所有可评估模型中,标记点均为最具破坏性的适配窗口;而在多项选择控制任务中则符号一致但衰减明显,符合生成长度加剧损伤的规律。该扫描提供了实用指导:无需训练的LRD预筛选和默认深层放置规则,尽管相对于无适配基线的绝对增益仍较小。最后,看似由表示稳定性损失带来的增益,在充足生成预算下逆转——截断链式思维被误判为空值,这对任何基于链式思维任务的干预评估提出方法论警示。

原文摘要 · Abstract (English)

When a language model fails on surface-perturbed input (typos, OCR noise, homophones), "which layer is responsible" has three natural operationalizations: where representations diverge most (sensitivity), where restoring clean activations recovers the prediction (causality), and where a small adapter can repair the damage (compensatory capacity) - and we show these three layer maps dissociate. Across a five-model panel we identify two propagation regimes - spike-and-suppress (Phi-3.5, Gemma-2-9B) and late-accumulation (Llama-3, Mistral, Qwen2.5-7B) - and on the two models meeting an 80% identity-patch gate, sensitivity and causality are anti-correlated (rho = -0.72 to -0.88). Within-family scaling on Qwen2.5 (1.5B to 14B) shows the late-accumulation signature strengthening monotonically with scale, corroborated on a second family. We propose cascade disruption as the mechanism behind the dissociation: adapters placed at causally implicated early layers break intact downstream computation, making diagnostic-flagged sites the worst adapter placements. A fixed-harness layer sweep across four models (3.8-8B) confirms the core prediction on chain-of-thought GSM8K - the flagged sites are the most damaging adapter windows on every adjudicable model - and is sign-consistent but strongly attenuated on a multiple-choice control, consistent with damage that compounds with generation length. The sweep yields practical guidance: a training-free LRD pre-screen and a default-deepest placement rule, though absolute gains over no-adapter baselines remain small. Finally, apparent gains from a representation-stability loss reverse under an adequate generation budget - truncated chain-of-thought had been scored as empty - a methodological warning for any intervention evaluated on chain-of-thought tasks.

模型诊断鲁棒性分析层间机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。