arXiv:2412.15320cs.CV2024-12被引 4

让模型同时抵御多种有害用途,提升开源安全性。

Multi-concept Model Immunization through Differentiable Model Merging

  • 通过可微分合并层,统一训练多概念免疫初始化。
  • 在多个概念上验证了免疫效果,优于单概念方法。
  • 适合关注模型安全与可控开源的研究者。

模型免疫是一种新兴方向,旨在降低开源模型被滥用的风险,通过使模型权重难以在特定有害任务上微调,实现‘免疫’。现有工作仅针对单一概念,但现实场景需抵御多种概念。为此,我们提出一种多概念免疫算法,通过可微分合并层,联合学习适应多种概念的统一‘难初始化’。实验中,我们将先前针对再学习和个性化适配的设置扩展至多概念场景,验证了该方法的有效性。

原文摘要 · Abstract (English)

Model immunization is an emerging direction that aims to mitigate the potential risk of misuse associated with open-sourced models and advancing adaptation methods. The idea is to make the released models' weights difficult to fine-tune on certain harmful applications, hence the name ``immunized''. Recent work on model immunization focuses on the single-concept setting. However, models need to be immunized against multiple concepts in real-world situations. To address this gap, we propose an immunization algorithm that, simultaneously, learns a single ``difficult initialization'' for adaptation methods over a set of concepts. We achieve this by incorporating a differentiable merging layer that combines a set of model weights adapted over multiple concepts. In our experiments, we demonstrate the effectiveness of multi-concept immunization by generalizing prior work's experiment setup of re-learning and personalization adaptation to multiple concepts.

模型安全多概念可微分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。