微调会意外破坏模型安全,因优化轨迹会自动滑向敏感区域。
The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety
- 发现安全方向在高维空间中呈低维曲面,易被梯度下降意外触发
- 提出四次方缩放定律:安全损失随训练时间的四次方增长
- 适合关注模型安全与微调机制的研究者阅读
在无害数据上对对齐语言模型进行微调时,安全防护机制会意外退化,即使开发者无恶意意图。现有解释认为微调更新应与安全关键方向正交,但该假设结构不稳定,会在梯度下降过程中崩溃。本文通过新几何分析证明:对齐集中在低维曲面上,具有显著曲率,导致一阶优化方法无法察觉或防御。尽管初始更新可能避开这些区域,但微调损失的曲率会产生二阶加速,系统性地将参数轨迹引向敏感区域。我们提出了对齐不稳定性条件,即三个几何特征共同满足时会导致安全退化。主要结果揭示了一个四次方缩放律:对齐损失随训练时间的四次方增长,受对齐几何尖锐性和微调任务与安全参数间曲率耦合强度的控制。这些发现暴露了当前安全范式中的结构性盲点。主流安全微调方法仅关注初始状态,忽略了根本性的动态过程。对齐脆弱性并非可修复的缺陷,而是梯度下降在弯曲流形上的内在几何属性。研究呼吁发展曲率感知方法,并推动对齐安全分析从被动红队测试转向面向开放权重模型部署的预测性诊断。
原文摘要 · Abstract (English)
Fine-tuning aligned language models on benign tasks unpredictably degrades safety guardrails, even when training data contains no harmful content and developers have no adversarial intent. We show that the prevailing explanation, that fine-tuning updates should be orthogonal to safety-critical directions in high-dimensional parameter space, offers false reassurance: we show this orthogonality is structurally unstable and collapses under the dynamics of gradient descent. We then resolve this through a novel geometric analysis, proving that alignment concentrates in low-dimensional subspaces with sharp curvature, creating a brittle structure that first-order methods cannot detect or defend. While initial fine-tuning updates may indeed avoid these subspaces, the curvature of the fine-tuning loss generates second-order acceleration that systematically steers trajectories into alignment-sensitive regions. We formalize this mechanism through the Alignment Instability Condition, three geometric properties that, when jointly satisfied, lead to safety degradation. Our main result establishes a quartic scaling law: alignment loss grows with the fourth power of training time, governed by the sharpness of alignment geometry and the strength of curvature coupling between the fine-tuning task and safety-critical parameters. These results expose a structural blind spot in the current safety paradigm. The dominant approaches to safe fine-tuning address only the initial snapshot of a fundamentally dynamic problem. Alignment fragility is not a bug to be patched; it is an intrinsic geometric property of gradient descent on curved manifolds. Our results motivate the development of curvature-aware methods, and we hope will further enable a shift in alignment safety analysis from reactive red-teaming to predictive diagnostics for open-weight model deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。