通过迭代优化上下文提升语义漂移越狱攻击效果
ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization

- 基于上下文语义漂移能力迭代优化攻击提示
- 在8个模型上平均攻击成功率达74.6%
- 适合研究模型安全与对抗样本的学者参考
基础模型在多种任务中表现卓越,但仍存在安全漏洞。语义漂移越狱攻击通过将有害词汇替换为无害替代词,并利用上下文诱导模型重新解读这些替代词为原有害概念,从而绕过安全机制。然而现有方法效果有限。本文发现,这一局限源于忽视上下文的语义漂移能力。系统分析表明,具备更强语义漂移能力的上下文更易引导模型恢复有害含义,实现成功越狱。基于此,我们识别并提炼有效上下文特征,提出黑盒上下文感知的语义漂移越狱框架ICO。每轮迭代中,ICO结合特征与目标模型反馈优化上下文。在三个数据集和八个基础模型上的实验表明,ICO持续优于八种先进基线,平均攻击成功率达74.6%。
原文摘要 · Abstract (English)
Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift jailbreaks have recently emerged as a promising attack paradigm. They bypass explicit safety mechanisms by replacing harmful terms in original harmful questions with benign alternatives and leveraging contextual information to induce the target model to reinterpret these alternatives as their corresponding harmful concepts. However, existing semantic-shift jailbreaks often achieve limited effectiveness. In this work, we reveal that this limitation arises from overlooking the semantic-shift capability of contexts. Through systematic analysis, we find that contexts exhibit substantially different abilities in inducing semantic shifts: contexts with stronger semantic-shift capabilities are more likely to guide models toward recovering harmful meanings and achieving successful jailbreaks. Based on this finding, we systematically identify and distill the characteristics of effective contexts and propose a black-box context-aware semantic-shift jailbreak framework with Iterative Context Optimization (ICO). In each iteration, ICO leverages these characteristics and feedback from the target model to optimize contexts. Extensive experiments on three datasets and eight target foundation models demonstrate that ICO consistently outperforms eight state-of-the-art baselines, achieving an average attack success rate of 74.6%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。