提出Anti-Diffusion系统,防护扩散模型被滥用的风险
Anti-Diffusion: Preventing Abuse of Modifications of Diffusion-Based Models
- 用提示词调优策略提升原图表达精度
- 引入语义扰动损失,有效防御编辑与调参攻击
- 构建Defense-Edit数据集,评估防御效果
尽管基于扩散的技术在图像生成与编辑任务中表现卓越,但其滥用可能引发严重的社会负面影响。近期一些工作尝试防御扩散模型的滥用,但其保护效果在特定场景下受限于人工定义提示词或稳定扩散(SD)版本,且仅关注调参方法,忽视了同样具有威胁性的编辑方法。为此,本文提出Anti-Diffusion——一种适用于通用扩散模型的隐私保护系统,兼顾调参与编辑技术的防护。为克服人工提示词对防御性能的限制,我们引入提示词调优(PT)策略,实现对原始图像的精准表达;为同时防御调参与编辑方法,提出语义扰动损失(SDL),破坏受保护图像的语义信息。鉴于针对编辑方法的防御研究有限,我们构建了名为Defense-Edit的数据集,用于评估各类方法的防御性能。实验表明,Anti-Diffusion在多种扩散模型与场景下均展现出优越的防御能力。
原文摘要 · Abstract (English)
Although diffusion-based techniques have shown remarkable success in image generation and editing tasks, their abuse can lead to severe negative social impacts. Recently, some works have been proposed to provide defense against the abuse of diffusion-based methods. However, their protection may be limited in specific scenarios by manually defined prompts or the stable diffusion (SD) version. Furthermore, these methods solely focus on tuning methods, overlooking editing methods that could also pose a significant threat. In this work, we propose Anti-Diffusion, a privacy protection system designed for general diffusion-based methods, applicable to both tuning and editing techniques. To mitigate the limitations of manually defined prompts on defense performance, we introduce the prompt tuning (PT) strategy that enables precise expression of original images. To provide defense against both tuning and editing methods, we propose the semantic disturbance loss (SDL) to disrupt the semantic information of protected images. Given the limited research on the defense against editing methods, we develop a dataset named Defense-Edit to assess the defense performance of various methods. Experiments demonstrate that our Anti-Diffusion achieves superior defense performance across a wide range of diffusion-based techniques in different scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。