arXiv:2602.07919cs.AIcs.CV2026-02被引 1

动态定位并精准剔除有害概念神经元,提升扩散模型安全性与效率

Selective Fine-Tuning for Targeted and Robust Concept Unlearning

  • 基于海森矩阵动态识别有害概念对应的神经元
  • 可同时清除单一、组合及条件性有害概念,且生成质量保持良好
  • 比现有方法更快更鲁棒,适用于实际部署的模型安全加固

文本引导的扩散模型被数百万用户使用,但容易被用于生成有害内容。概念剔除方法旨在降低模型生成有害内容的可能性。传统方法仅针对单一概念,少数近期工作考虑了更现实的概念组合。然而,当前最先进方法依赖全量微调,计算成本高。概念定位方法可支持选择性微调,但现有技术为静态方案,效果不佳。为此,我们提出 TRUST(目标化稳健选择性微调),一种通过海森矩阵正则化动态估计目标概念神经元并进行选择性微调的新方法。实验表明,相较于多个 SOTA 基线,TRUST 对对抗提示具有强鲁棒性,显著保留生成质量,且速度远超现有方法。该方法无需特定正则化即可实现单一、组合及条件概念的有效剔除。

原文摘要 · Abstract (English)

Text guided diffusion models are used by millions of users, but can be easily exploited to produce harmful content. Concept unlearning methods aim at reducing the models' likelihood of generating harmful content. Traditionally, this has been tackled at an individual concept level, with only a handful of recent works considering more realistic concept combinations. However, state of the art methods depend on full finetuning, which is computationally expensive. Concept localisation methods can facilitate selective finetuning, but existing techniques are static, resulting in suboptimal utility. In order to tackle these challenges, we propose TRUST (Targeted Robust Selective fine Tuning), a novel approach for dynamically estimating target concept neurons and unlearning them through selective finetuning, empowered by a Hessian based regularization. We show experimentally, against a number of SOTA baselines, that TRUST is robust against adversarial prompts, preserves generation quality to a significant degree, and is also significantly faster than the SOTA. Our method achieves unlearning of not only individual concepts but also combinations of concepts and conditional concepts, without any specific regularization.

扩散模型概念剔除模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。