arXiv:2602.20880cs.CV2026-02被引 11

动态识别有害类别冲突,避免安全引导互相干扰。

When Safety Collides: Resolving Multi-Category Harmful Conflicts in Text-to-Image Diffusion via Adaptive Safety Guidance

  • 根据生成状态动态识别最相关的有害类别
  • 仅对识别出的类别施加安全引导,减少多类冲突
  • 无需训练,可适配不同安全防护场景

文本到图像扩散模型在生成高质量图像方面取得显著进展,但也带来潜在的有害内容生成风险。现有基于安全引导的方法通过平均多个有害类别关键词来规避有害区域,但无法捕捉不同危害类别间的复杂相互作用,导致‘有害冲突’——即缓解一类危害可能无意中加剧另一类,反而提高整体有害率。为此,我们提出冲突感知自适应安全引导(CASG),一种无需训练的框架,在生成过程中动态识别并应用与类别对齐的安全方向。CASG由两部分组成:(i) 冲突感知类别识别(CaCI),用于识别与模型当前生成状态最匹配的有害类别;(ii) 冲突化解引导应用(CrGA),仅沿识别出的类别施加安全引导,避免多类别干扰。CASG可应用于潜空间和文本空间的安全防护。在T2I安全基准上的实验表明,相比现有方法,其最高可将有害率降低15.4%,达到当前最佳性能。

原文摘要 · Abstract (English)

Text-to-Image (T2I) diffusion models have demonstrated significant advancements in generating high-quality images, while raising potential safety concerns regarding harmful content generation. Safety-guidance-based methods have been proposed to mitigate harmful outputs by steering generation away from harmful zones, where the zones are averaged across multiple harmful categories based on predefined keywords. However, these approaches fail to capture the complex interplay among different harm categories, leading to "harmful conflicts" where mitigating one type of harm may inadvertently amplify another, thus increasing overall harmful rate. To address this issue, we propose Conflict-aware Adaptive Safety Guidance (CASG), a training-free framework that dynamically identifies and applies the category-aligned safety direction during generation. CASG is composed of two components: (i) Conflict-aware Category Identification (CaCI), which identifies the harmful category most aligned with the model's evolving generative state, and (ii) Conflict-resolving Guidance Application (CrGA), which applies safety steering solely along the identified category to avoid multi-category interference. CASG can be applied to both latent-space and text-space safeguards. Experiments on T2I safety benchmarks demonstrate CASG's state-of-the-art performance, reducing the harmful rate by up to 15.4% compared to existing methods.

安全生成扩散模型对抗性引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。