arXiv:2602.07340cs.LG2026-02被引 4

通过控制参数空间几何结构,提升大模型安全对齐的鲁棒性。

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control

  • 基于优化几何视角,设计选择性参数子空间控制策略
  • 在多种噪声偏好设置下,安全对齐鲁棒性显著优于主流方法
  • 适用于存在分布偏移的场景,适合模型安全增强研究者

大语言模型的安全对齐在领域偏移和噪声偏好监督下仍显脆弱。现有鲁棒对齐方法多关注对齐数据的不确定性,却忽视了基于偏好的目标函数所引发的优化过程中的脆弱性。本文从优化几何角度重新审视这一问题,认为仅靠数据驱动方法无法解决鲁棒性失效。提出ShaPO框架,通过在对齐关键参数子空间上实施选择性几何控制,强制最坏情况下的对齐目标。相比全局约束,该方法避免了过度正则化带来的分布偏移下性能下降。我们在两个层级实现:词元级ShaPO稳定似然替代优化,奖励级ShaPO在噪声监督下保持奖励一致性。在多个安全基准和噪声偏好设置下,ShaPO持续优于主流偏好优化方法。且能与数据鲁棒目标良好组合,进一步提升性能,验证了优化几何视角的有效性。代码已开源。

原文摘要 · Abstract (English)

Safety alignment of large language models remains brittle under domain shift and noisy preference supervision. Most existing robust alignment methods focus on uncertainty in alignment data, while overlooking optimization-induced fragility in preference-based objectives. In this work, we revisit robustness for LLM safety alignment from an optimization geometry perspective, and argue that robustness failures cannot be addressed by data-centric methods alone. We propose \textit{ShaPO}, a geometry-aware preference optimization framework that enforces worst-case alignment objectives via selective geometry control over alignment-critical parameter subspace. By avoiding uniform geometry constraints, ShaPO mitigates the over-regularization that can harm robustness under distribution shift. We instantiate ShaPO at two levels: token-level ShaPO stabilizes likelihood-based surrogate optimization, while reward-level ShaPO enforces reward-consistent optimization under noisy supervision. Across diverse safety benchmarks and noisy preference settings, ShaPO consistently improves safety robustness over popular preference optimization methods. Moreover, ShaPO composes cleanly with data-robust objectives, yielding additional gains and empirically supporting the proposed optimization-geometry perspective. The code is available at https://github.com/liujilong0116/ShaPO.

大模型安全对齐优化鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。