arXiv:2510.01223cs.CRcs.CL2025-10被引 2

用语义相关的嵌套场景隐藏攻击意图,突破大模型安全防线

Jailbreaking LLMs via Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge

  • 构建与查询语义高度相关且含特定有害知识的嵌套场景
  • 在GPT-4o、Llama3-70b等模型上实现高效越狱,隐蔽性强
  • 适合研究模型安全防御机制或红队测试的人员使用

大型语言模型(LLMs)在各类任务中表现卓越,但仍易受越狱攻击,引发有害输出。嵌套场景策略虽具潜力,但因明显恶意意图易被检测。本文首次发现并系统验证:当嵌套场景与查询高度语义相关并融合针对性有害知识时,模型对这类场景不敏感,该方向尚未被充分探索。基于此,提出RTS-Attack框架,通过构建语义相关且含特定有害知识的场景,实现对大模型对齐机制的自适应自动化绕过。生成的越狱提示不含直接有害查询,隐蔽性优异。大量实验表明,该方法在效率和通用性上均优于基线,在GPT-4o、Llama3-70b、Gemini-pro等先进模型上表现突出。完整代码已公开。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities in various tasks. However, they remain exposed to jailbreak attacks, eliciting harmful responses. The nested scenario strategy has been increasingly adopted across various methods, demonstrating immense potential. Nevertheless, these methods are easily detectable due to their prominent malicious intentions. In this work, we are the first to find and systematically verify that LLMs' alignment defenses are not sensitive to nested scenarios, where these scenarios are highly semantically relevant to the queries and incorporate targeted toxic knowledge. This is a crucial yet insufficiently explored direction. Based on this, we propose RTS-Attack (Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge), an adaptive and automated framework to examine LLMs' alignment. By building scenarios highly relevant to the queries and integrating targeted toxic knowledge, RTS-Attack bypasses the alignment defenses of LLMs. Moreover, the jailbreak prompts generated by RTS-Attack are free from harmful queries, leading to outstanding concealment. Extensive experiments demonstrate that RTS-Attack exhibits superior performance in both efficiency and universality compared to the baselines across diverse advanced LLMs, including GPT-4o, Llama3-70b, and Gemini-pro. Our complete code is available at https://github.com/nercode/Work. WARNING: THIS PAPER CONTAINS POTENTIALLY HARMFUL CONTENT.

模型安全越狱攻击语义相关

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。