跨大模型迁移安全干预,用激活空间映射实现高效对齐
Activation Space Interventions Can Be Transferred Between Large Language Models
- 通过共享激活空间映射,实现安全指令在不同大模型间的迁移
- 在回溯攻击清除与有害请求拒答任务中成功转移控制向量
- 适用于模型对齐、安全开关等场景,尤其适合资源受限环境
AI模型表征通用性的研究揭示了跨领域、多模态与架构的日益趋同。然而,表征通用性的实际应用仍待探索。本文通过学习模型间共享激活空间的映射,首次证明安全干预可在不同大模型间有效迁移。我们在两个经典安全任务上验证:回溯攻击移除与有害提示拒答,成功实现了可预测输出调控的控制向量迁移。此外,我们提出新任务“受损能力”,即对模型进行微调以嵌入与后门关联的知识,测试其分离有用技能与后门的能力,反映真实世界挑战。在Llama、Qwen和Gemma系列模型上的广泛实验表明,该方法可利用小模型高效对齐大模型。进一步发现,基座模型与微调模型间的自编码器映射可作为可靠的“轻量级安全开关”,支持动态行为切换。
原文摘要 · Abstract (English)
The study of representation universality in AI models reveals growing convergence across domains, modalities, and architectures. However, the practical applications of representation universality remain largely unexplored. We bridge this gap by demonstrating that safety interventions can be transferred between models through learned mappings of their shared activation spaces. We demonstrate this approach on two well-established AI safety tasks: backdoor removal and refusal of harmful prompts, showing successful transfer of steering vectors that alter the models' outputs in a predictable way. Additionally, we propose a new task, \textit{corrupted capabilities}, where models are fine-tuned to embed knowledge tied to a backdoor. This tests their ability to separate useful skills from backdoors, reflecting real-world challenges. Extensive experiments across Llama, Qwen and Gemma model families show that our method enables using smaller models to efficiently align larger ones. Furthermore, we demonstrate that autoencoder mappings between base and fine-tuned models can serve as reliable ``lightweight safety switches", allowing dynamic toggling between model behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。