arXiv:2410.15396cs.CRcs.AI2024-10被引 6

利用LLM自身缺陷,反制其驱动的自动化网络攻击。

The Best Defense is a Good Offense: Countering LLM-Powered Cyberattacks

  • 从LLM的偏见、信任输入等弱点入手设计防御
  • 黑盒测试下最高90%成功率,可有效阻断攻击
  • 适合安全研究者与防御系统开发者参考

随着大语言模型(LLMs)不断发展,其被用于自动化网络攻击的可能性日益增加。凭借侦察、漏洞利用和命令执行等能力,LLMs可能成为自主网络攻击代理的核心。本文提出新型防御策略,利用攻击性LLM的内在缺陷。通过针对其偏见、对输入的信任、记忆限制以及单一思维模式等弱点,我们开发出误导、延迟或中和这些自主代理的技术。在黑盒条件下评估防御效果,从单次提示-响应场景逐步扩展至使用自建CTF机器的真实测试。实验结果表明,防御成功率最高达90%,证明将LLM自身漏洞转化为防御手段具有显著有效性。

原文摘要 · Abstract (English)

As large language models (LLMs) continue to evolve, their potential use in automating cyberattacks becomes increasingly likely. With capabilities such as reconnaissance, exploitation, and command execution, LLMs could soon become integral to autonomous cyber agents, capable of launching highly sophisticated attacks. In this paper, we introduce novel defense strategies that exploit the inherent vulnerabilities of attacking LLMs. By targeting weaknesses such as biases, trust in input, memory limitations, and their tunnel-vision approach to problem-solving, we develop techniques to mislead, delay, or neutralize these autonomous agents. We evaluate our defenses under black-box conditions, starting with single prompt-response scenarios and progressing to real-world tests using custom-built CTF machines. Our results show defense success rates of up to 90\%, demonstrating the effectiveness of turning LLM vulnerabilities into defensive strategies against LLM-driven cyber threats.

网络安全LLM防御智能攻防

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。