arXiv:2501.18626cs.CRcs.AI2025-01ACL被引 7

用提示词藏任务破解大模型安全限制,6个主流模型均被攻破

The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs

  • 在提示词中嵌入解密、谜题等序列任务,诱导模型生成违规内容
  • 在6个顶尖模型上成功绕过防护,包括GPT-4o和LLaMA 3.2
  • 揭示当前对齐机制漏洞,适合安全研究者关注

我们提出一类新型越狱对抗攻击——任务在提示(Task-in-Prompt, TIP)攻击。该方法将序列到序列任务(如密码解码、谜题解答、代码执行)嵌入模型提示中,间接诱导生成禁止内容。为系统评估攻击效果,我们构建了PHRYGE基准测试。实验表明,该技术可成功绕过六个前沿语言模型的安全防护,涵盖GPT-4o与LLaMA 3.2。研究揭示了当前大模型安全对齐中的关键缺陷,凸显亟需更先进的防御策略。警告:本文包含伦理不当提问示例,仅用于学术研究。

原文摘要 · Abstract (English)

We present a novel class of jailbreak adversarial attacks on LLMs, termed Task-in-Prompt (TIP) attacks. Our approach embeds sequence-to-sequence tasks (e.g., cipher decoding, riddles, code execution) into the model's prompt to indirectly generate prohibited inputs. To systematically assess the effectiveness of these attacks, we introduce the PHRYGE benchmark. We demonstrate that our techniques successfully circumvent safeguards in six state-of-the-art language models, including GPT-4o and LLaMA 3.2. Our findings highlight critical weaknesses in current LLM safety alignments and underscore the urgent need for more sophisticated defence strategies. Warning: this paper contains examples of unethical inquiries used solely for research purposes.

越狱攻击大模型安全对抗样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。