arXiv:2510.02194cs.AIcs.CR2025-10

通过动态调控模型安全与实用性的平衡,提升大模型的安全可控性。

UpSafe$^\circ$C: Upcycling for Controllable Safety in Large Language Models

  • 将关键安全层重构为稀疏专家混合结构,用路由器实现软性安全控制。
  • 在多个基准上显著降低有害内容生成,同时保持通用能力竞争力。
  • 支持推理时灵活调节安全温度,实现安全与性能的帕累托最优。

大型语言模型在众多任务中取得显著进展,但仍面临生成有害内容和越狱攻击等安全风险。现有安全技术——包括外部防护、推理阶段引导和后训练对齐——在安全、实用性与可控性之间难以平衡。本文提出UpSafe$^ ext{°C}$,一种通过安全感知的模型复用(upcycling)增强大模型安全性的统一框架。该方法首先识别安全关键层,并将其重构为稀疏的专家混合(MoE)结构,其中路由器作为软性防护机制,选择性激活原始MLP与新增的安全专家。进一步引入两阶段SFT策略,在强化安全判别能力的同时保留通用能力。为实现推理时的灵活控制,提出安全温度机制,可动态调整安全与实用性的权衡。多基准、多基础模型及不同规模下的实验表明,UpSafe$^ ext{°C}$ 在对抗有害输入和越狱攻击方面实现稳健的安全提升,同时在通用任务上保持竞争力。分析显示,安全温度提供细粒度的推理期控制,达到安全与性能的帕累托最优前沿。结果揭示了大模型安全的新方向:从静态对齐转向动态、模块化、推理感知的控制。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved remarkable progress across a wide range of tasks, but remain vulnerable to safety risks such as harmful content generation and jailbreak attacks. Existing safety techniques -- including external guardrails, inference-time guidance, and post-training alignment -- each face limitations in balancing safety, utility, and controllability. In this work, we propose UpSafe$^\circ$C, a unified framework for enhancing LLM safety through safety-aware upcycling. Our approach first identifies safety-critical layers and upcycles them into a sparse Mixture-of-Experts (MoE) structure, where the router acts as a soft guardrail that selectively activates original MLPs and added safety experts. We further introduce a two-stage SFT strategy to strengthen safety discrimination while preserving general capabilities. To enable flexible control at inference time, we introduce a safety temperature mechanism, allowing dynamic adjustment of the trade-off between safety and utility. Experiments across multiple benchmarks, base model, and model scales demonstrate that UpSafe$^\circ$C achieves robust safety improvements against harmful and jailbreak inputs, while maintaining competitive performance on general tasks. Moreover, analysis shows that safety temperature provides fine-grained inference-time control that achieves the Pareto-optimal frontier between utility and safety. Our results highlight a new direction for LLM safety: moving from static alignment toward dynamic, modular, and inference-aware control.

大模型安全专家混合推理控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。