攻击者用隐性文字注入,让文生图模型生成带恶意文本的图像。
Reading Between the Pixels: An Inscriptive Jailbreak Attack on Text-to-Image Models
- 将提示词拆成语义伪装、视觉定位、字体编码三层,逐层优化攻击
- 在7个模型上平均成功率65.57%,最高达91.00%
- 揭示现有安全机制对文字渲染的防御盲点,适合安全研究者参考
现代文生图(T2I)模型已能生成可读的段落级文字,催生了一类新型滥用行为。本文提出并形式化了‘隐性越狱’攻击:攻击者诱导T2I系统在视觉无害场景中嵌入有害文本内容(如伪造文件)。与传统以视觉违规为目标的越狱不同,此类攻击利用文本生成能力本身作为武器。现有方法针对粗粒度视觉操控设计,难以在保持字符级精度的同时绕过多阶段安全过滤。为此,我们提出Etch——一个黑盒攻击框架,将对抗性提示分解为语义伪装、视觉-空间锚定和排版编码三个功能正交层,将全提示空间联合优化转化为可解的子问题,并通过零阶迭代循环与视觉语言模型协同优化。该模型评估7个T2I模型、2个基准数据集,平均攻击成功率达65.57%(峰值91.00%),显著优于现有基线。结果揭示当前T2I安全对齐存在关键盲区,亟需具备排版感知能力的多模态防御机制。
原文摘要 · Abstract (English)
Modern text-to-image (T2I) models can now render legible, paragraph-length text, enabling a fundamentally new class of misuse. We identify and formalize the inscriptive jailbreak, where an adversary coerces a T2I system into generating images containing harmful textual payloads (e.g., fraudulent documents) embedded within visually benign scenes. Unlike traditional depictive jailbreaks that elicit visually objectionable imagery, inscriptive attacks weaponize the text-rendering capability itself. Because existing jailbreak techniques are designed for coarse visual manipulation, they struggle to bypass multi-stage safety filters while maintaining character-level fidelity. To expose this vulnerability, we propose Etch, a black-box attack framework that decomposes the adversarial prompt into three functionally orthogonal layers: semantic camouflage, visual-spatial anchoring, and typographic encoding. This decomposition reduces joint optimization over the full prompt space to tractable sub-problems, which are iteratively refined through a zero-order loop. In this process, a vision-language model critiques each generated image, localizes failures to specific layers, and prescribes targeted revisions. Extensive evaluations across 7 models on the 2 benchmarks demonstrate that Etch achieves an average attack success rate of 65.57% (peaking at 91.00%), significantly outperforming existing baselines. Our results reveal a critical blind spot in current T2I safety alignments and underscore the urgent need for typography-aware defense multimodal mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。