用条件化非线性传输,让AI生成更安全图像同时不丢质量。
Conditioned Activation Transport for T2I Safety Steering

- 基于几何条件机制,只在危险区域调整激活值。
- 在2300组近似提示对上测试,攻击成功率显著下降。
- 适合需要内容安全的文本生成场景,如广告、社交平台。
当前文本到图像(T2I)模型仍易生成不当内容。尽管激活调制可在推理时干预,但线性调制常导致良性提示下图像质量下降。为此,我们构建了包含2300对高余弦相似度的安全与不安全提示的对比数据集SafeSteerDataset。基于此,提出条件激活传输(CAT)框架,采用基于几何的条件机制与非线性传输映射。通过仅在不安全激活区域触发传输,最大限度减少对良性查询的干扰。在Z-Image和Infinity两个先进架构上验证,CAT在保持图像保真度的同时,显著降低攻击成功率,具备良好泛化能力。警告:本文含潜在冒犯性文字与图像。
原文摘要 · Abstract (English)
Despite their impressive capabilities, current Text-to-Image (T2I) models remain prone to generating unsafe and toxic content. While activation steering offers a promising inference-time intervention, we observe that linear activation steering frequently degrades image quality when applied to benign prompts. To address this trade-off, we first construct SafeSteerDataset, a contrastive dataset containing 2300 safe and unsafe prompt pairs with high cosine similarity. Leveraging this data, we propose Conditioned Activation Transport (CAT), a framework that employs a geometry-based conditioning mechanism and nonlinear transport maps. By conditioning transport maps to activate only within unsafe activation regions, we minimize interference with benign queries. We validate our approach on two state-of-the-art architectures: Z-Image and Infinity. Experiments demonstrate that CAT generalizes effectively across these backbones, significantly reducing Attack Success Rate while maintaining image fidelity compared to unsteered generations. Warning: This paper contains potentially offensive text and images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。