arXiv:2603.22041cs.CV2026-03被引 1

提出双阶段干预框架,有效防止文本生成有害图像。

DTVI: Dual-Stage Textual and Visual Intervention for Safe Text-to-Image Generation

  • 分两阶段干预:先在文本全序列识别恶意语义,再在图像生成时削弱残留风险。
  • 在7类有害内容上平均防御成功率88.56%,性相关类别达94.43%。
  • 适合关注生成安全、需对抗恶意提示的研究者与应用开发者。

文本到图像(T2I)扩散模型生成能力强大,但可能产生不安全内容,引发严重安全问题。现有推理阶段防御方法通常在文本嵌入空间进行无类别区分的词元级干预,难以捕捉分布于完整词元序列中的恶意语义,且易受对抗性提示攻击。本文提出DTV I,一种双阶段推理阶段防御框架,用于安全T2I生成。不同于仅干预特定词元嵌入的方法,本方法在完整提示嵌入上实施类别感知的序列级干预,以更好捕捉分散的恶意语义,并在视觉生成阶段进一步抑制剩余不安全影响。在真实世界不安全提示、对抗性提示及多个有害类别上的实验表明,该方法在保持良性提示生成质量的同时,实现有效且鲁棒的防御,跨性类别基准平均防御成功率(DSR)达94.43%,跨七类不安全内容达88.56%。

原文摘要 · Abstract (English)

Text-to-Image (T2I) diffusion models have demonstrated strong generation ability, but their potential to generate unsafe content raises significant safety concerns. Existing inference-time defense methods typically perform category-agnostic token-level intervention in the text embedding space, which fails to capture malicious semantics distributed across the full token sequence and remains vulnerable to adversarial prompts. In this paper, we propose DTVI, a dual-stage inference-time defense framework for safe T2I generation. Unlike existing methods that intervene on specific token embeddings, our method introduces category-aware sequence-level intervention on the full prompt embedding to better capture distributed malicious semantics, and further attenuates the remaining unsafe influences during the visual generation stage. Experimental results on real-world unsafe prompts, adversarial prompts, and multiple harmful categories show that our method achieves effective and robust defense while preserving reasonable generation quality on benign prompts, obtaining an average Defense Success Rate (DSR) of 94.43% across sexual-category benchmarks and 88.56 across seven unsafe categories, while maintaining generation quality on benign prompts.

文本生成安全防御扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。