构建真实网络安全攻防评估框架,测试大模型实际渗透能力。
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
- 设计四类真实场景,模拟单点、混合、链式和带防御的漏洞利用。
- 7个前沿大模型均无法突破防御,复杂场景下表现普遍不佳。
- 适合安全研究者、模型开发者,用于评估AI在攻防中的可信性。
大型语言模型(LLMs)自主性的提升亟需对其潜在网络攻击能力进行严格评估。现有基准普遍缺乏真实世界复杂性,难以准确衡量LLMs的网络安全能力。为此,我们提出PACEbench,一个基于真实漏洞难度、环境复杂性和网络安全防御原则的实用型人工智能网络攻防评估框架。该框架包含四种场景:单点漏洞利用、混合漏洞利用、链式漏洞利用及带防御的漏洞利用。为应对这些复杂挑战,我们提出PACEagent,一种模拟人类渗透测试人员的新型智能体,支持多阶段侦察、分析与攻击。对七个前沿大模型的广泛实验表明,当前模型在复杂网络场景中表现有限,且均未能绕过防御机制。结果表明,现有模型尚未构成通用性网络攻击威胁。本工作为未来模型的可信发展提供了可靠评估工具。
原文摘要 · Abstract (English)
The increasing autonomy of Large Language Models (LLMs) necessitates a rigorous evaluation of their potential to aid in cyber offense. Existing benchmarks often lack real-world complexity and are thus unable to accurately assess LLMs' cybersecurity capabilities. To address this gap, we introduce PACEbench, a practical AI cyber-exploitation benchmark built on the principles of realistic vulnerability difficulty, environmental complexity, and cyber defenses. Specifically, PACEbench comprises four scenarios spanning single, blended, chained, and defense vulnerability exploitations. To handle these complex challenges, we propose PACEagent, a novel agent that emulates human penetration testers by supporting multi-phase reconnaissance, analysis, and exploitation. Extensive experiments with seven frontier LLMs demonstrate that current models struggle with complex cyber scenarios, and none can bypass defenses. These findings suggest that current models do not yet pose a generalized cyber offense threat. Nonetheless, our work provides a robust benchmark to guide the trustworthy development of future models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。