提出新方法与评估指标,让图像抵御恶意编辑更有效。
Semantic Mismatch and Perceptual Degradation: A New Perspective on Image Editing Immunity
- 通过操控扩散模型中间特征,同时破坏语义对齐和引发视觉劣化。
- 在多个测试中实现最高免疫成功率,显著优于现有方法。
- 适合关注数字内容安全、对抗生成攻击的研究者使用。
基于文本引导的扩散模型图像编辑虽强大,但存在被滥用风险,促使研究者尝试通过不可察觉的扰动保护图像免受未经授权的修改。现有评估方法通常依赖输出与参考图像的视觉差异,但这忽略了图像免疫的核心目标——即破坏编辑结果与攻击者意图之间的语义一致性,无论其是否偏离特定输出。本文主张,免疫成功应定义为编辑结果要么与提示语义不匹配,要么出现显著感知劣化,从而挫败恶意目的。为此,提出协同中间特征操纵(SIFM)方法,通过双重协同目标:(1) 最大化中间特征与原始编辑轨迹的差异,以破坏语义对齐;(2) 最小化特征范数,诱导感知劣化。同时引入免疫成功率(ISR)这一新指标,首次严格量化真实免疫效果,利用多模态大语言模型评估语义失败或显著感知劣化的比例。大量实验表明,SIFM在防御扩散模型恶意操作方面达到当前最优性能。
原文摘要 · Abstract (English)
Text-guided image editing via diffusion models, while powerful, raises significant concerns about misuse, motivating efforts to immunize images against unauthorized edits using imperceptible perturbations. Prevailing metrics for evaluating immunization success typically rely on measuring the visual dissimilarity between the output generated from a protected image and a reference output generated from the unprotected original. This approach fundamentally overlooks the core requirement of image immunization, which is to disrupt semantic alignment with attacker intent, regardless of deviation from any specific output. We argue that immunization success should instead be defined by the edited output either semantically mismatching the prompt or suffering substantial perceptual degradations, both of which thwart malicious intent. To operationalize this principle, we propose Synergistic Intermediate Feature Manipulation (SIFM), a method that strategically perturbs intermediate diffusion features through dual synergistic objectives: (1) maximizing feature divergence from the original edit trajectory to disrupt semantic alignment with the expected edit, and (2) minimizing feature norms to induce perceptual degradations. Furthermore, we introduce the Immunization Success Rate (ISR), a novel metric designed to rigorously quantify true immunization efficacy for the first time. ISR quantifies the proportion of edits where immunization induces either semantic failure relative to the prompt or significant perceptual degradations, assessed via Multimodal Large Language Models (MLLMs). Extensive experiments show our SIFM achieves the state-of-the-art performance for safeguarding visual content against malicious diffusion-based manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。