arXiv:2604.15967cs.CRcs.CV2026-04

多个安全概念组合竟生成有害图像,揭示文本生成模型的隐性安全漏洞。

When Safe Concepts Become Unsafe: Multi-Concept Compositional Vulnerabilities in Text-to-Image Models

论文配图:When Safe Concepts Become Unsafe: Multi-Concept Compositional Vulnerabilities in Text-to-Image Models
图 1 · 摘自论文原文
  • 通过组合多个安全概念触发模型生成有害内容,无需恶意提示。
  • 13个模型中,FLUX.1在99.35%情况下产生不安全输出,防御率仅1.57%。
  • 现有防护方法对未知组合无效,暴露生成模型的组合推理风险。

文本到图像(T2I)模型越来越擅长忠实遵循用户指令,但我们发现这种能力引出了名为多概念组合不安全(MCCU)的安全隐患。当多个单独安全的概念在单个生成请求中组合时,会生成有害或敏感的视觉内容。与以往需要对抗性提示或模型访问的越狱方式不同,MCCU不依赖恶意输入,而是源于模型对多个安全视觉概念的组合生成机制。为系统评估该风险,我们构建了TwoHamsters框架,包含20,000条提示、51组精选概念对和6类风险类别。我们在黑盒设置下评估了13个T2I模型。结果表明,指令遵循能力与安全性存在明显冲突:越能忠实遵循提示的模型,越容易发生MCCU失败。例如,FLUX.1的不安全对齐率达99.35%,但仅1.57%的防御率。我们进一步测试了三种典型防御策略——安全过滤、MCCU专用检测器微调和概念擦除,均无法抵御未见过的概念组合。研究暗示,当前安全机制未能覆盖生成模型在组合推理中的攻击面。我们计划公开TwoHamsters,推动社区发展更先进的生成防御技术。

原文摘要 · Abstract (English)

Text-to-image (T2I) models are increasingly optimized for following user instructions faithfully. However, we find that this capability introduces a safety vulnerability we call Multi-Concept Compositional Unsafety (MCCU). MCCU occurs when multiple individually safe concepts, if combined in a single generation request, lead to harmful or sensitive visual outputs. Unlike prior jailbreak settings, MCCU does not rely on adversarial prompts, model access, or explicitly disallowed content. Instead, the risk emerges from how the model composes multiple safe visual concepts into a single scene. To systematically measure this threat, we build TwoHamsters, a large-scale evaluation framework consisting of 20k prompts, 51 curated concept pairs, and six risk categories. We evaluate 13 T2I models under a black-box setting. Our results show a clear conflict between instruction-following and safety: models that follow prompts more faithfully tend to produce more MCCU failures. For example, FLUX.1 achieves a 99.35% Unsafe Alignment Rate while only reaching a 1.57% MCCU Defense Rate. We further evaluate three representative defenses, including safety filtering, MCCU-specific detector fine-tuning, and concept erasure, all of which fail against unseen concept combinations. Our findings suggest that compositional reasoning in T2I models creates an attack surface that is not captured by existing safety mechanisms. We anticipate the release of TwoHamsters will catalyze community development of advanced generative defense mechanisms.

安全漏洞文本生成组合推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。