提出新型提示攻击方法,揭示扩散模型文本编码器的脆弱性
CAHS-Attack: CLIP-Aware Heuristic Search Attack Method for Stable Diffusion
- 结合蒙特卡洛树搜索与遗传算法,高效生成对抗性提示
- 在短长提示上均实现当前最佳攻击效果
- 揭示基于CLIP的文生图系统存在根本性安全风险
扩散模型在对抗性提示下表现出显著脆弱性,强化攻击能力对发现漏洞和构建更鲁棒生成系统至关重要。现有方法多依赖白盒模型梯度或手工提示工程,在实际部署中因访问受限或攻击效果差而不可行。本文提出CAHS-Attack,一种面向CLIP的启发式搜索攻击方法:利用约束遗传算法预筛选高潜力对抗提示作为根节点,结合蒙特卡洛树搜索进行细粒度后缀优化,并在每次模拟回溯中保留语义破坏力最强的结果以实现高效局部搜索。大量实验表明,该方法在不同语义的短提示和长提示上均达到当前最优攻击性能。此外,我们发现扩散模型的脆弱性源于其基于CLIP的文本编码器固有缺陷,揭示了当前文生图流水线存在的根本性安全风险。
原文摘要 · Abstract (English)
Diffusion models exhibit notable fragility when faced with adversarial prompts, and strengthening attack capabilities is crucial for uncovering such vulnerabilities and building more robust generative systems. Existing works often rely on white-box access to model gradients or hand-crafted prompt engineering, which is infeasible in real-world deployments due to restricted access or poor attack effect. In this paper, we propose CAHS-Attack , a CLIP-Aware Heuristic Search attack method. CAHS-Attack integrates Monte Carlo Tree Search (MCTS) to perform fine-grained suffix optimization, leveraging a constrained genetic algorithm to preselect high-potential adversarial prompts as root nodes, and retaining the most semantically disruptive outcome at each simulation rollout for efficient local search. Extensive experiments demonstrate that our method achieves state-of-the-art attack performance across both short and long prompts of varying semantics. Furthermore, we find that the fragility of SD models can be attributed to the inherent vulnerability of their CLIP-based text encoders, suggesting a fundamental security risk in current text-to-image pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。