测试大模型自主渗透能力,发现部分模型已能无指导攻破服务器
The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems

- 构建双层靶机环境+通用代理框架,模拟真实攻防场景
- 19个模型渗透成功率10.7%至69.3%,模型越强越能自主攻击
- 揭示大模型在网络安全中的潜在风险,适合关注AI安全的研究者
当前,自主执行可能造成重大现实危害的网络攻击,被视为前沿AI系统不可逾越的关键红线。在这一背景下,自主渗透作为核心使能能力,指大模型驱动的AI系统无需人工干预,独立对目标服务器开展对抗性操作、识别并利用漏洞、获取未经授权访问或控制的能力。现有评估常采用不透明方法、脱离实际的简化场景,或给予大模型过多先验知识与任务特定指导,难以真实反映现代AI系统在高影响力网络攻击场景下的自主渗透能力。为此,我们构建了一个新的自主渗透评估框架,包含两部分:靶机服务器与代理支撑架构。靶机侧设计两级环境,基于部署的无已知漏洞服务数量区分:一级(1个安全服务)与二级(3个安全服务),共生成300个靶机。代理支撑采用通用代理架构,配备通用网络安全工具,不提供任何目标特异性先验知识。评估了19个开源与专有大模型,结果显示当前模型渗透成功率在10.7%至69.3%之间。此外,观察到自主渗透能力随模型整体能力提升而持续增强。
原文摘要 · Abstract (English)
Nowadays, the autonomous execution of cyberattacks capable of causing substantial real-world harm is widely regarded as one of the critical red lines that frontier AI systems must not cross. Within this broader red-line scenario, autonomous penetration represents a core enabling capability and subtask: the ability of LLM-powered AI systems to independently conduct adversarial operations against a target server without human intervention, identify and exploit vulnerabilities, and obtain unauthorized access or control. A growing body of work has sought to assess the autonomous penetration capabilities of AI systems. However, existing evaluations often employ opaque methodologies, rely on unrealistic or overly simplified penetration-testing scenarios, or provide LLMs with excessive prior knowledge and task-specific guidance, and cannot accurately capture the extent to which modern AI systems can autonomously perform this core capability within broader high-impact cyberattack scenarios. To address these limitations, we construct a new autonomous penetration evaluation framework consisting of two components: target servers and agent scaffolding. Specifically, on the target-server side, we design two levels of target environments based on the number of secure services without known vulnerabilities deployed alongside a vulnerable service: Tier~1 (one secure service) and Tier~2 (three secure services), resulting in a total of 300 target servers. Meanwhile, the agent scaffolding adopts a general-purpose agent architecture equipped with a set of general-purpose cybersecurity tools, without any target-specific prior knowledge. We evaluate 19 open-weight and proprietary LLMs, and find that current models achieve penetration success rates ranging from 10.7% to 69.3%. Moreover, we observe that autonomous penetration capability continues to improve alongside advances in overall model capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。