arXiv:2607.27594cs.LG2026-07

用超网络生成LoRA适配器,实现大模型按需安全对齐。

Compliance2LoRA: Personalizable On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters

论文配图:Compliance2LoRA: Personalizable On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters
图 1 · 摘自论文原文
  • 通过超网络动态生成适配不同安全策略的LoRA权重。
  • 单个大模型可支持多种策略组合,无需重复训练。
  • 适合需要个性化安全控制的下游应用部署。

大型推理模型(LRMs)的后训练对齐显著提升了其在多样化安全合规场景中的适应能力。然而,随着模型个性化需求增长,不同用户所需的合规策略子集差异增大,为每个策略子集单独训练模型带来严重的组合爆炸问题。现有上下文学习方法虽缓解了这一问题,但引入了长上下文生成的计算负担。为此,我们提出 extsc{Compliance2LoRA},一种基于统一自适应超网络的多策略合规框架。该框架将安全策略作为可定制输入,由超网络生成对应的LoRA适配器权重,注入原始大模型后即可生成符合指定策略子集的响应。实验表明,该方法可在不牺牲不同规模模型及多评估数据集上任务性能的前提下,实现对单一模型的按需策略调整,验证了基于超网络的自适应对齐在大型推理模型中的有效性与实用性。

原文摘要 · Abstract (English)

Post-training alignment in large reasoning models (LRMs) has significantly improved their adaptability to diverse safety compliance settings. However, as LRMs personalization for downstream users takes center stage, the demand for varying levels of policy compliance grows as different user-specific LRMs must adhere to distinct subsets of safety policies. Training a separate LRM for each policy subset introduces severe combinatorial overhead. While in context learning methods overcome this combinatorial overhead, they introduce additional computational challenges associated with long context generation. To address this challenge, we propose \ours, a unified adaptive hypernetwork-based framework for multi-policy compliance. In our framework, safety policies serve as customizable inputs to a LoRA adapter generator, which learns to produce policy compliant LoRA weights for downstream LRM. When added to the LRM these weights enable the generation of responses compliant with the specified policy subsets. In this work, we demonstrate that training such a hypernetwork enables on-demand policy adjustments on a single LRM without sacrificing task performance across reasoning models of different sized and different evaluation datasets. This highlights the effectiveness and practicality of adaptive hypernetwork based alignment in LRMs.

安全对齐LoRA超网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。