分析大模型在各种输入扰动下的表现,揭示单看输出会误导对模型鲁棒性的判断。
How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models

- 从输出、隐藏层几何、注意力机制三层次分析六类扰动传播
- 不同扰动导致可区分的表征变化,且跨模型一致性有限
- 建议用多层级评估替代单一指标,适合研究模型稳健性者
语言模型常遭遇拼写错误、文本损坏、词汇替换和词元顺序打乱,但鲁棒性通常仅通过输出行为评估。本文在六种自然与合成扰动下,对四个GPT-2和两个Qwen2.5检查点进行三层分析:输出行为、隐藏状态几何与注意力头功能。采用中心核对齐与内在维度分析逐层几何,考察GPT-2的注意力响应。不同扰动产生可区分的度量特征,这些特征未被输出指标完全捕捉,且在各检查点间部分不一致。复制得分与词元替换及重排下的激活修复密切相关。梯度引导的HotFlip扰动在GPT-2中引发的行为与表征破坏强于等速率随机替换,其行为效应在全部六个检查点中保持一致。结果表明,仅依赖单一行为或表征指标的鲁棒性结论可能具有误导性,亟需多层次评估扰动对语言模型计算过程的影响。
原文摘要 · Abstract (English)
Language models encounter typos, corrupted text, altered words, and disrupted token order, yet robustness is usually evaluated only through output behavior. We study how six naturalistic and synthetic input perturbations propagate through decoder-only language models at three levels: output behavior, hidden-state geometry, and attention-head function. We evaluate behavioral effects across four GPT-2 and two Qwen2.5 checkpoints by analyzing layerwise geometry using centered kernel alignment and intrinsic dimension, and examine attention-head responses in GPT-2. Perturbation types produce distinguishable metric profiles that are not fully captured by output measures and are only partly consistent across the tested checkpoints. Copying scores are especially associated with activation-patching recovery under token substitution and shuffling. Gradient-guided HotFlip perturbations also cause stronger behavioral and representational disruption than rate-matched random token substitutions in GPT-2; their behavioral effects are consistent across all six tested checkpoints. Our results show that robustness claims based on a single behavioral or representational metric can be misleading, and motivate multi-level evaluation of how perturbations alter language-model computation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。