arXiv:2410.08776cs.CRcs.AI2024-10被引 3

黑客伪造安全检测结果,骗过大模型生成有害内容

F2A: An Innovative Approach for Prompt Injection by Utilizing Feign Security Detection Agents

  • 用伪造的安全检测结果欺骗大模型的防御机制
  • 实验显示多个主流大模型均被成功劫持
  • 适合关注大模型安全与对抗攻击的研究者

随着大语言模型(LLMs)的快速发展,其在内容安全检测领域已广泛应用。然而我们发现,大模型对安全检测代理存在盲目信任。攻击者可利用此漏洞,通过在提示中注入虚假安全检测结果,绕过防御机制,获取有害内容并劫持正常对话。本文提出一种名为伪装代理攻击(Feign Agent Attack, F2A)的新攻击方法。通过一系列实验,验证了F2A在多种场景下对主流大模型的劫持能力,并分析了其根本原因——大模型缺乏对增强代理输出的批判性评估。论文还提出了相应解决方案,强调大模型应主动质疑外部检测结果,以显著提升系统可靠性与安全性,防范此类攻击。

原文摘要 · Abstract (English)

With the rapid development of Large Language Models (LLMs), numerous mature applications of LLMs have emerged in the field of content safety detection. However, we have found that LLMs exhibit blind trust in safety detection agents. The general LLMs can be compromised by hackers with this vulnerability. Hence, this paper proposed an attack named Feign Agent Attack (F2A).Through such malicious forgery methods, adding fake safety detection results into the prompt, the defense mechanism of LLMs can be bypassed, thereby obtaining harmful content and hijacking the normal conversation. Continually, a series of experiments were conducted. In these experiments, the hijacking capability of F2A on LLMs was analyzed and demonstrated, exploring the fundamental reasons why LLMs blindly trust safety detection results. The experiments involved various scenarios where fake safety detection results were injected into prompts, and the responses were closely monitored to understand the extent of the vulnerability. Also, this paper provided a reasonable solution to this attack, emphasizing that it is important for LLMs to critically evaluate the results of augmented agents to prevent the generating harmful content. By doing so, the reliability and security can be significantly improved, protecting the LLMs from F2A.

大模型安全对抗攻击提示注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。