arXiv:2603.13028cs.CRcs.AI2026-03被引 1

提出新方法破解图像保护,实现一次净化、自由编辑。

Purify Once, Edit Freely: Breaking Image Protections under Model Mismatch

  • 用隐空间投影和指令引导重建,无须知道原图或防御机制
  • 在2100次编辑任务中恢复可编辑性,提升3-6 dB PSNR
  • 揭示保护一旦被破就失效,提醒需考虑模型差异风险

扩散模型虽能实现高质量图像编辑,但也可能被滥用进行未经授权的风格模仿和有害内容生成。为缓解风险,主动防护方法在分享前向图像嵌入微小且通常不可察觉的对抗扰动,以破坏后续编辑或微调。然而,在真实发布后场景中,内容所有者无法控制下游处理流程,针对代理模型优化的防护在攻击者使用不匹配的扩散管道时可能失效。现有净化方法虽能削弱防护,但常牺牲图像质量,且很少考察架构差异。本文提出统一的发布后净化框架,评估防护在模型不匹配下的生存能力。设计两种实用净化器:VAE-Trans通过隐空间投影修正受保护图像;EditorClean利用扩散变换器进行指令引导重建,挖掘架构异质性。两者均无需访问受保护图像或防御内部信息。在2100次编辑任务和六种代表性防护方法上,EditorClean持续恢复可编辑性。相比受保护输入,其提升PSNR达3-6 dB,FID降低50-70%,优于先前基线约2 dB PSNR和30%更低的FID。结果揭示‘净化一次,自由编辑’的失败模式:一旦净化成功,保护信号基本消失,导致无限制编辑。这凸显了在模型不匹配下评估防护必要性,并需设计对异构攻击者鲁棒的防御。

原文摘要 · Abstract (English)

Diffusion models enable high-fidelity image editing but can also be misused for unauthorized style imitation and harmful content generation. To mitigate these risks, proactive image protection methods embed small, often imperceptible adversarial perturbations into images before sharing to disrupt downstream editing or fine-tuning. However, in realistic post-release scenarios, content owners cannot control downstream processing pipelines, and protections optimized for a surrogate model may fail when attackers use mismatched diffusion pipelines. Existing purification methods can weaken protections but often sacrifice image quality and rarely examine architectural mismatch. We introduce a unified post-release purification framework to evaluate protection survivability under model mismatch. We propose two practical purifiers: VAE-Trans, which corrects protected images via latent-space projection, and EditorClean, which performs instruction-guided reconstruction with a Diffusion Transformer to exploit architectural heterogeneity. Both operate without access to protected images or defense internals. Across 2,100 editing tasks and six representative protection methods, EditorClean consistently restores editability. Compared to protected inputs, it improves PSNR by 3-6 dB and reduces FID by 50-70 percent on downstream edits, while outperforming prior purification baselines by about 2 dB PSNR and 30 percent lower FID. Our results reveal a purify-once, edit-freely failure mode: once purification succeeds, the protective signal is largely removed, enabling unrestricted editing. This highlights the need to evaluate protections under model mismatch and design defenses robust to heterogeneous attackers.

图像安全扩散模型净化攻击模型差异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。