揭露文本生成图像模型安全对齐中的虚假高可用性陷阱
The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models

- 用细粒度评估发现安全对齐后语义忠实度严重下降
- 提出SAGE方法,恢复嵌入空间结构,提升细粒度可用性5.0%
- 适合关注生成质量与安全平衡的研究者和开发者
文本到图像扩散模型的安全对齐旨在抑制有害生成的同时保留良性提示下的可用性。现有方法看似兼具高安全与高可用性,实则依赖粗粒度全局指标(如FID、CLIPScore),这些指标对细粒度语义正确性不敏感,造成虚假的高可用性印象。我们通过结构化评估TIFA(基于问答的文本到图像忠实度评估)发现,安全对齐模型在物体数量、属性和关系等维度上出现显著语义失真。分析提示嵌入空间揭示了语义坍缩现象:嵌入分布收缩且提示间相似性结构扭曲,与结构化可用性损失强相关。基于此,我们提出结构感知几何正则化(SAGE),在对齐过程中显式保持嵌入分散性与提示间关系结构。SAGE在保持强安全性能的同时,使结构化可用性相比先前最优提升5.0%,粗粒度指标亦具竞争力。代码与模型已开源。
原文摘要 · Abstract (English)
Safety alignment of text-to-image (T2I) diffusion models aims to suppress harmful generations while preserving utility on benign prompts. Recent methods often appear to deliver high safety with high utility, but this conclusion rests largely on coarse global utility metrics (e.g., FID, CLIPScore) that are insensitive to fine-grained semantic correctness, creating an illusion of high utility. We show that when utility is measured with structured evaluation, this illusion breaks: on TIFA (Text-to-Image Faithfulness evaluation with Question Answering), safety-aligned models suffer substantial drops in semantic fidelity, including failures in object counts, attributes, and relationships. To diagnose the source of this gap, we analyze the text-encoder prompt embedding space and uncover semantic collapse, a contraction of embedding spread coupled with distortion of inter-prompt similarity structure, which strongly correlates with structured utility loss. Guided by this insight, we propose StructureAware Geometric Regularization (SAGE), a safety alignment objective that explicitly preserves embedding spread and inter-prompt relational structure during adaptation. Our method restores structured utility (TIFA +5.0% over prior state-of-the-art) while maintaining strong safety performance and competitive coarse-grained utility scores. Our source code and trained models are available at https://adeelyousaf.github.io/SAGE_ECCV26_Project_Page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。