arXiv:2505.17646cs.LG2025-05被引 12

发现大模型损失曲面中存在稳定区域,可解释其抗干扰能力与微调鲁棒性。

Unveiling the Basin-Like Loss Landscape in Large Language Models

  • 通过分析参数空间扰动,发现大模型存在性能稳定的'盆地'区域。
  • 预训练形成基础能力盆地,微调构建特定能力盆地(如安全、数学)。
  • 建议微调限定在盆地内以保留原有能力,适合关注模型稳定性研究者。

我们发现大型语言模型的损失曲面中存在'盆地'现象。随着模型规模增大,模型对参数空间中的随机扰动表现出更强的鲁棒性,形成广阔且性能相近的稳定区域,而超出该区域时模型能力会急剧下降。预训练构建了'基础能力盆地',后续对齐微调则形成'特定能力盆地'(如安全性、数学、编程)。因此我们认为,限制微调在盆地内的良性微调可保留先前能力。此外,我们分析了最坏方向的损失曲面,发现其始终尖锐且有害。对抗性微调沿近似最坏方向进行,导致模型能力快速退化。最后,理论分析表明,盆地大小约束了任何微调(包括对抗性微调)的性能退化程度,同时保障了模型对输入扰动的鲁棒性,提示扩大盆地具有实际益处。

原文摘要 · Abstract (English)

We discover the emergence of \textit{basins} in the loss landscape of large language models. As model scale increases, LLMs become progressively more resilient to random perturbations in the parameter space, giving rise to expansive stability regions where models exhibit nearly identical performance, but outside of which their capabilities collapse. We observe that pre-training creates a \textit{basic capability} basin, and subsequent alignment fine-tuning forms \textit{specific capability} basins (e.g., safety, math, coding). Thus, we argue that benign fine-tuning confined to the basin should preserve prior capabilities. Besides, we also analyze the loss landscape for worst-case directions, which is consistently sharp and detrimental. We find that adversarial fine-tuning moves along the nearly worst-case directions, thus rapidly degrading model capabilities. Finally, we provide a theoretical analysis demonstrating that the basin size bounds the performance degradation of any fine-tuning, including the adversarial ones, while also guaranteeing the model robustness w.r.t. input perturbations, suggesting the benefit of enlarging basins.

大模型损失曲面微调鲁棒性盆地现象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。