用工具伪装攻击大模型,让有害内容绕过安全检测。
Jailbreaking Large Language Models through Iterative Tool-Disguised Attacks via Reinforcement Learning
- 将恶意请求伪装成正常工具调用,绕过内容过滤。
- 通过多轮对话逐步提升回复危害性,攻击成功率更高。
- 适合研究模型安全与对抗攻击的人员参考。
大型语言模型(LLMs)在诸多应用中表现出卓越能力,但仍极易受到越狱攻击,导致生成违背人类价值观和安全准则的有害内容。尽管已有大量防御研究,现有防护机制仍难以应对复杂的对抗策略。本文提出iMIST(交互式多步渐进工具伪装越狱攻击),一种新型自适应越狱方法,协同利用当前防御机制的漏洞。iMIST将恶意查询伪装为正常工具调用以规避内容过滤,并引入交互式渐进优化算法,通过多轮对话结合实时危害性评估,动态提升回复危害程度。实验表明,iMIST在多个主流模型上均实现了更高的攻击有效性,同时保持较低的拒绝率。结果揭示了当前大模型安全机制的关键缺陷,凸显了构建更鲁棒防御策略的紧迫性。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated remarkable capabilities across diverse applications, however, they remain critically vulnerable to jailbreak attacks that elicit harmful responses violating human values and safety guidelines. Despite extensive research on defense mechanisms, existing safeguards prove insufficient against sophisticated adversarial strategies. In this work, we propose iMIST (\underline{i}nteractive \underline{M}ulti-step \underline{P}rogre\underline{s}sive \underline{T}ool-disguised Jailbreak Attack), a novel adaptive jailbreak method that synergistically exploits vulnerabilities in current defense mechanisms. iMIST disguises malicious queries as normal tool invocations to bypass content filters, while simultaneously introducing an interactive progressive optimization algorithm that dynamically escalates response harmfulness through multi-turn dialogues guided by real-time harmfulness assessment. Our experiments on widely-used models demonstrate that iMIST achieves higher attack effectiveness, while maintaining low rejection rates. These results reveal critical vulnerabilities in current LLM safety mechanisms and underscore the urgent need for more robust defense strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。