开源大模型可独立发动网络攻击,需警惕其安全威胁。
Countering Autonomous Cyber Threats
- 用开源大模型测试在隔离网络中攻陷目标机的能力。
- 最新可下载模型与领先专有模型在攻击效果上相当。
- 通过恶意提示注入可有效阻断AI攻击流程,适合安全研究者。
基础模型具备生成自然语言和代码的能力,带来双重用途风险,尤其在网络安全领域。生成式AI已通过数百个恶意AI即服务工具影响网络空间,协助开发恶意软件与社会工程攻击。更令人担忧的是,近期研究表明这些先进模型可能自主执行或指导进攻性网络行动。然而,以往研究多聚焦专有模型,因此前缺乏强开源权重模型,且未探索网络防御或应对措施的影响。鉴于可下载模型更难监管和防止滥用,评估其作为进攻性网络代理的潜力至关重要。本研究评估了多个前沿基础模型在孤立网络中攻陷商用目标机的能力,并检验防御机制的有效性。结果表明,最新发布的可下载模型在使用常见黑客工具针对已知漏洞时,性能与领先专有模型相当。为缓解此类大模型驱动的威胁,研究展示防御性提示注入(DPI)payload可有效破坏恶意攻击者的操作流程。由此结果,分析了人工智能安全与治理在网络安全中的深远影响。
原文摘要 · Abstract (English)
With the capability to write convincing and fluent natural language and generate code, Foundation Models present dual-use concerns broadly and within the cyber domain specifically. Generative AI has already begun to impact cyberspace through a broad illicit marketplace for assisting malware development and social engineering attacks through hundreds of malicious-AI-as-a-services tools. More alarming is that recent research has shown the potential for these advanced models to inform or independently execute offensive cyberspace operations. However, these previous investigations primarily focused on the threats posed by proprietary models due to the until recent lack of strong open-weight model and additionally leave the impacts of network defenses or potential countermeasures unexplored. Critically, understanding the aptitude of downloadable models to function as offensive cyber agents is vital given that they are far more difficult to govern and prevent their misuse. As such, this work evaluates several state-of-the-art FMs on their ability to compromise machines in an isolated network and investigates defensive mechanisms to defeat such AI-powered attacks. Using target machines from a commercial provider, the most recently released downloadable models are found to be on par with a leading proprietary model at conducting simple cyber attacks with common hacking tools against known vulnerabilities. To mitigate such LLM-powered threats, defensive prompt injection (DPI) payloads for disrupting the malicious cyber agent's workflow are demonstrated to be effective. From these results, the implications for AI safety and governance with respect to cybersecurity is analyzed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。