arXiv:2508.09190cs.LGcs.AI2025-08

不重训练即可提升微调大模型的安全性,防止越狱攻击。

Multi-Level Safety Continual Projection for Fine-Tuned Large Language Models without Retraining

  • 通过多层级表征协同,隐式对齐全局与局部安全激活。
  • 在极小参数扰动下,有害输出减少且任务性能不变。
  • 可持续防御新出现的安全风险,适合部署后加固场景。

微调服务虽快速扩展大语言模型的任务能力,但常导致安全对齐表征退化与重组,使模型更易偏离人类偏好并暴露于新型越狱攻击。现有后微调防御方法多依赖单尺度安全修正,难以兼顾安全性、模型效用与持续适应性。本文提出无需重训练的多层级安全持续投影(MSCP)方法,通过协同多层级表征隐式对齐全局与局部安全激活,隔离控制敏感行为的稀疏神经元簇,并实施可组合的安全方向投影,有效抑制有害输出,仅需微小参数扰动即可保持任务性能并提升与人类偏好的对齐。大量实验表明,该方法显著降低有害性评分与攻击成功率,同时保留模型效用。此外,引入任务特定的多维异构安全激活聚类机制,实现对未知新兴安全威胁的持续防御与泛化能力。

原文摘要 · Abstract (English)

While fine-tuning services drive the rapid expansion of task capabilities in large language models (LLMs), they are often accompanied by the degradation and reorganization of safety-aligned representations, making models more prone to deviating from human preferences and exposing them to emerging jailbreak risks. Existing post-fine-tuning defense methods predominantly rely on single-scale safety correction mechanisms, which struggle to achieve a robust balance among safety, model utility, and continual adaptability. We propose Multi-Level Safety Continual Projection (MSCP), a training-free post-fine-tuning safety enhancement method that implicitly aligns global and localized safety activations through coordinated multi-level representations to isolate sparse neuron clusters governing safety-sensitive behaviors. It then applies composable safety-direction projections without retraining, effectively suppressing harmful outputs under minimal parameter perturbations while preserving task performance and improving alignment with human preferences. Extensive experiments across multiple fine-tuned LLM models demonstrate that our method significantly reduce harmfulness scores and attack success rates with minimal parameter modifications, while preserving the model's utility. Furthermore, we introduce a task-specific, multi-dimensional heterogeneous safety activation clustering mechanism that enables continual defense and generalization capability against unforeseen emerging safety concerns.

大模型安全微调加固无重训练越狱防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。