提出新方法在文本生成图像中绕过多重安全检测,提升攻击成功率。
Token-Level Constraint Boundary Search for Jailbreaking Text-to-Image Models
- 基于决策边界搜索近似最优提示词,优化生成过程
- 在全链路防御下实现52.5%的攻击成功率达新高
- 适用于研究模型安全或对抗样本的开发者
近年来,文本到图像(T2I)生成技术快速发展,但其生成有害内容的安全风险也日益突出。实际部署中,T2I服务通常采用包含提示词检查器、安全训练生成器和后处理图像检查器的全链路防御机制。在黑盒环境下,攻击此类系统极具挑战性,因为提示词构成离散组合空间,且攻击需在稀疏反馈和有限查询条件下满足多重耦合约束。为此,本文提出一种新型基于查询的黑盒越狱攻击——TCBS-Attack,通过在文本与图像检查器定义的决策边界附近搜索提示词,将决策边界作为约束条件引导提示词种群的进化搜索,迭代优化接近边界的提示词。该方法有效缩小搜索空间,提升查询效率并保持语义连贯性。大量实验表明,TCBS-Attack在多种T2I模型上持续优于现有最先进越狱攻击,包括安全训练的开源模型及DALL-E 3等商业在线服务。在全链路防御下,其平均攻击成功率ASR-4达52.5%,ASR-1为22.0%,显著超越基线方法。
原文摘要 · Abstract (English)
Text-to-Image (T2I) generation has advanced rapidly in recent years, but they also raise safety concerns due to the potential production of harmful content. In the practical deployments, T2I services typically adopt full-chain defenses that combine a prompt checker, a securely trained generator, and a post-hoc image checker. Jailbreaking such full-chain systems is challenging in the black-box settings because prompt tokens form a discrete combinatorial space and the attack must satisfy multiple coupled constraints under sparse feedback and limited queries. To address these challenges, we propose Token-level Constraint Boundary Search (TCBS)-Attack, a novel query-based black-box jailbreak attack that searches for tokens located near the decision boundaries defined by text and image checkers. TCBS-Attack incorporates decision boundaries as constraint conditions to guide the evolutionary search of token populations, iteratively optimize tokens near these boundaries. Such evolutionary search process reduces the effective search space and improves query efficiency while preserving semantic coherence. Extensive experiments demonstrate that TCBS-Attack consistently outperforms state-of-the-art jailbreak attacks across various T2I models, including securely trained open-source models and commercial online services like DALL-E 3. TCBS-Attack achieves an ASR-4 of 52.5% and an ASR-1 of 22.0% on jailbreaking full-chain T2I models, significantly surpassing baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。