arXiv:2506.00382cs.LGcs.CL2025-06ACL被引 3

发现大模型中不依赖数据的内在关键层,提升可解释性与鲁棒性。

Spectral Insights into Data-Oblivious Critical Layers in Large Language Models

  • 通过中心核对齐分析预训练模型层间表征动态,识别关键层。
  • 关键层在微调中变化最大,且其表征由前几主成分驱动。
  • 适用于高效领域适应和对抗攻击防御,提升性能与安全性。

理解大语言模型中特征表示随层数演进的机制,是提升模型可解释性与鲁棒性的关键。现有研究虽已识别出与特定功能相关的关键层,但多依赖微调后的数据相关分析,仅适用于事后分析。本文提出一种数据无关的方法,在未微调的预训练模型中,通过中心核对齐(CKA)分析表征动态,识别内在关键层。我们发现,表征空间发生显著变化的层,正是微调时最敏感的层,这一规律在不同任务间保持一致。谱分析进一步揭示,这些变化由顶层主成分的改变驱动,编码从推理到结论的语义转变。我们将此发现应用于两个实际场景:在高效领域适应中,仅微调关键层即可实现比非关键层更大的损失下降;在后门防御中,冻结关键层可使攻击成功率降低高达40%。

原文摘要 · Abstract (English)

Understanding how feature representations evolve across layers in large language models (LLMs) is key to improving their interpretability and robustness. While recent studies have identified critical layers linked to specific functions or behaviors, these efforts typically rely on data-dependent analyses of fine-tuned models, limiting their use to post-hoc settings. In contrast, we introduce a data-oblivious approach to identify intrinsic critical layers in pre-fine-tuned LLMs by analyzing representation dynamics via Centered Kernel Alignment(CKA). We show that layers with significant shifts in representation space are also those most affected during fine-tuning--a pattern that holds consistently across tasks for a given model. Our spectral analysis further reveals that these shifts are driven by changes in the top principal components, which encode semantic transitions from rationales to conclusions. We further apply these findings to two practical scenarios: efficient domain adaptation, where fine-tuning critical layers leads to greater loss reduction compared to non-critical layers; and backdoor defense, where freezing them reduces attack success rates by up to 40%.

大模型表征分析微调优化安全防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。