提出黑盒攻击框架,揭示图像生成模型去记忆的脆弱性
REFORGE: Multi-modal Attacks Reveal Vulnerable Concept Unlearning in Image Generation Models
- 用笔画初始化图像,通过注意力引导掩码分配噪声
- 攻击成功率显著提升,同时保持语义一致性和视觉质量
- 适合研究生成模型安全与鲁棒性的人参考
近期图像生成模型(IGMs)在高保真内容生成方面取得进展,但也带来版权再现和生成不当内容等风险。图像生成模型去记忆(IGMU)通过移除有害概念来缓解这些风险,而无需完整重训练。尽管关注度上升,但在对抗性输入,尤其是黑盒设置下的图像侧威胁方面的鲁棒性仍缺乏研究。为此,我们提出REFORGE,一种基于对抗性图像提示的黑盒红队评估框架。REFORGE 初始化笔画图像,并采用跨注意力引导的掩码策略,将噪声分配至概念相关区域,平衡攻击效果与视觉保真度。在多个代表性去记忆任务与防御方法上的实验表明,REFORGE 显著提升攻击成功率,同时实现更强的语义对齐与更高效率,优于现有基线。结果揭示了当前IGMU方法存在的持续脆弱性,强调需针对多模态对抗攻击发展具备鲁棒性的去记忆机制。代码已开源:https://github.com/Imfatnoily/REFORGE。
原文摘要 · Abstract (English)
Recent progress in image generation models (IGMs) enables high-fidelity content creation but also amplifies risks, including the reproduction of copyrighted content and the generation of offensive content. Image Generation Model Unlearning (IGMU) mitigates these risks by removing harmful concepts without full retraining. Despite growing attention, the robustness under adversarial inputs, particularly image-side threats in black-box settings, remains underexplored. To bridge this gap, we present REFORGE, a black-box red-teaming framework that evaluates IGMU robustness via adversarial image prompts. REFORGE initializes stroke-based images and optimizes perturbations with a cross-attention-guided masking strategy that allocates noise to concept-relevant regions, balancing attack efficacy and visual fidelity. Extensive experiments across representative unlearning tasks and defenses demonstrate that REFORGE significantly improves attack success rate while achieving stronger semantic alignment and higher efficiency than involved baselines. These results expose persistent vulnerabilities in current IGMU methods and highlight the need for robustness-aware unlearning against multi-modal adversarial attacks. Our code is at: https://github.com/Imfatnoily/REFORGE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。