arXiv:2503.00038cs.CLcs.AI2025-03ACL被引 19

用隐喻诱导模型从善意内容生成有害回复,突破安全防护

from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors

  • 通过逻辑相关隐喻作为种子,诱导模型自我演化出有害内容
  • 在多个主流大模型上实现顶尖攻击成功率,具备良好迁移性
  • 揭示隐喻式攻击新路径,适合安全研究者关注

现有研究揭示了大语言模型(LLMs)因越狱攻击产生有害内容的风险,但忽略了直接生成有害内容比将良性内容转化为有害形式更困难。本研究提出一种新型攻击框架——自适应隐喻越狱(AVATAR),利用对抗性隐喻诱导模型将原本无害的内容校准为有害输出。具体而言,面对有害请求时,AVATAR 自适应识别一组与之逻辑相关但表面无害的隐喻作为初始种子;随后,借助这些隐喻,目标模型被引导进行推理与内容校准,从而实现越狱——或直接输出有害响应,或通过隐喻与专业有害内容间的残差差异完成校准。实验表明,AVATAR 能有效且可迁移地越狱多种先进大模型,在多个测试中达到当前最优攻击成功率。

原文摘要 · Abstract (English)

Current studies have exposed the risk of Large Language Models (LLMs) generating harmful content by jailbreak attacks. However, they overlook that the direct generation of harmful content from scratch is more difficult than inducing LLM to calibrate benign content into harmful forms. In our study, we introduce a novel attack framework that exploits AdVersArial meTAphoR (AVATAR) to induce the LLM to calibrate malicious metaphors for jailbreaking. Specifically, to answer harmful queries, AVATAR adaptively identifies a set of benign but logically related metaphors as the initial seed. Then, driven by these metaphors, the target LLM is induced to reason and calibrate about the metaphorical content, thus jailbroken by either directly outputting harmful responses or calibrating residuals between metaphorical and professional harmful content. Experimental results demonstrate that AVATAR can effectively and transferable jailbreak LLMs and achieve a state-of-the-art attack success rate across multiple advanced LLMs.

越狱攻击隐喻诱导安全风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。