arXiv:2604.03941cs.CV2026-04中稿 · 2026 IEEE Internat…

精准定位并消除图像危险区域,兼顾安全与细节保真

SafeCtrl: Region-Aware Safety Control for Text-to-Image Diffusion via Detect-Then-Suppress

  • 先检测后抑制:用注意力机制定位风险区域
  • 仅处理局部区域,生成内容保真度更高
  • 抗对抗攻击,适合需要可控生成的场景

文本到图像扩散模型的大规模部署面临生成视觉有害内容(如色情、暴力、恐怖图像)的严峻挑战。现有安全干预方法(从输入过滤到模型概念擦除)普遍存在两大缺陷:(1)安全与上下文保真之间存在严重权衡,去除不安全概念会损害安全内容的保真度;(2)易受对抗攻击,安全机制可被轻易绕过。为此,我们提出SafeCtrl,一种基于“检测-抑制”范式的区域感知安全控制框架。不同于全局干预,SafeCtrl首先通过注意力引导的检测模块精确定位特定风险区域;随后,基于图像级直接偏好优化(DPO)训练的局部抑制模块,仅在检测区域内中和有害语义,将不安全物体转化为安全替代物,同时完整保留周围上下文。跨多个风险类别的大量实验表明,SafeCtrl相比当前最优方法实现了更优的安全性与保真度权衡。关键的是,该方法对对抗提示攻击展现出更强鲁棒性,提供了一种精准且可靠的负责任生成解决方案。

原文摘要 · Abstract (English)

The widespread deployment of text-to-image diffusion models is significantly challenged by the generation of visually harmful content, such as sexually explicit content, violence, and horror imagery. Common safety interventions, ranging from input filtering to model concept erasure, often suffer from two critical limitations: (1) a severe trade-off between safety and context preservation, where removing unsafe concepts degrades the fidelity of the safe content, and (2) vulnerability to adversarial attacks, where safety mechanisms are easily bypassed. To address these challenges, we propose SafeCtrl, a Region-Aware safety control framework operating on a Detect-Then-Suppress paradigm. Unlike global safety interventions, SafeCtrl first employs an attention-guided Detect module to precisely localize specific risk regions. Subsequently, a localized Suppress module, optimized via image-level Direct Preference Optimization (DPO), neutralizes harmful semantics only within the detected areas, effectively transforming unsafe objects into safe alternatives while leaving the surrounding context intact. Extensive experiments across multiple risk categories demonstrate that SafeCtrl achieves a superior trade-off between safety and fidelity compared to state-of-the-art methods. Crucially, our approach exhibits improved resilience against adversarial prompt attacks, offering a precise and robust solution for responsible generation.

图像安全扩散模型对抗防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。