arXiv:2508.21099cs.CVcs.AI2025-08ACL被引 2

用外部模块修复生成图像安全问题,不降低画质。

Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification

  • 引入独立安全模块SafePatch,外置修正不改动原模型。
  • 在I2P数据集上仅7%内容不安全,保持图像质量。
  • 适合需要高安全性的图像生成应用,如内容审核。

文本到图像(T2I)生成模型虽具备出色视觉保真度,但仍易生成不安全内容。现有安全防御多在生成模型内部干预,但因概念纠缠严重,导致良性生成质量下降,这一代价被称为“安全税”。为克服此局限,本文倡导从破坏性内部编辑转向外部安全修正。基于此理念,提出SafePatch:一个结构隔离的安全模块,通过外部可解释修正实现安全防护,不修改基础模型。其核心为可训练的基模型编码器克隆,继承丰富语义先验并保持表征一致性。为实现可解释修正,构建严格对齐的反事实安全数据集(ACS),用于差异监督训练。在裸露及多类别基准测试中,以及针对近期对抗提示攻击,SafePatch在保持图像质量和语义一致性的前提下,实现鲁棒的不安全内容抑制,在I2P数据集上仅7%生成内容不安全。

原文摘要 · Abstract (English)

Text-to-image (T2I) generative models have achieved remarkable visual fidelity, yet remain vulnerable to generating unsafe content. Existing safety defenses typically intervene internally within the generative model, but suffer from severe concept entanglement, leading to degradation of benign generation quality, a trade-off we term the Safety Tax. To overcome this limitation, we advocate a paradigm shift from destructive internal editing to external safety rectification. Following this principle, we propose SafePatch, a structurally isolated safety module that performs external, interpretable rectification without modifying the base model. The core backbone of SafePatch is architecturally instantiated as a trainable clone of the base model's encoder, allowing it to inherit rich semantic priors and maintain representation consistency. To enable interpretable safety rectification, we construct a strictly aligned counterfactual safety dataset (ACS) for differential supervision training. Across nudity and multi-category benchmarks and recent adversarial prompt attacks, SafePatch achieves robust unsafe suppression (7% unsafe on I2P) while preserving image quality and semantic alignment.

图像生成安全检测外部修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。