arXiv:2410.21939cs.CRcs.AI2024-10被引 2

评测AI在真实软件中自动挖漏洞的能力,揭示其攻防风险。

HonestCyberEval: An AI Cyber Risk Benchmark for Automated Software Exploitation

  • 用带人工漏洞的Nginx代码库测试大模型自动化攻破能力。
  • o1-preview成功率最高达92.85%,o3-mini等更省钱但效果差。
  • 为评估AI网络攻击风险提供可复现的真实场景基准。

我们提出HonestCyberEval,一个用于评估AI模型在自动化软件漏洞利用中的能力与风险的新基准,重点关注其在真实软件系统中发现并利用漏洞的能力。评估基于增强合成漏洞的Nginx Web服务器仓库进行,涵盖多个主流语言模型,包括OpenAI的GPT-4.5、o3-mini、o1和o1-mini,Anthropic的Claude-3-7-sonnet-20250219、Claude-3.5-sonnet-20241022及Claude-3.5-sonnet-20240620,Google DeepMind的Gemini-1.5-pro,以及OpenAI的早期GPT-4o模型。结果显示,各模型在成功率与效率上差异显著,其中o1-preview达到最高成功率92.85%,而o3-mini和Claude-3.7-sonnet-20250219则提供了成本更低但成功率较低的替代方案。该风险评估为系统性衡量真实网络攻击场景下的AI安全风险奠定了基础。

原文摘要 · Abstract (English)

We introduce HonestCyberEval, a new benchmark for assessing AI models' capabilities and risks in automated software exploitation, focusing on their ability to detect and exploit vulnerabilities in real-world software systems. Our evaluation leverages the Nginx web server repository augmented with synthetic vulnerabilities. We assess several leading language models, including OpenAI's GPT-4.5, o3-mini, o1 and o1-mini, Anthropic's Claude-3-7-sonnet-20250219, Claude-3.5-sonnet-20241022 and Claude-3.5-sonnet-20240620, Google DeepMind's Gemini-1.5-pro, and OpenAI's earlier GPT-4o model. Our findings reveal that these models vary significantly in their success rates and efficiency, with o1-preview achieving the highest success rate (92.85\%) and o3-mini and Claude-3.7-sonnet-20250219 providing cost-effective but less successful alternatives. This risk evaluation establishes a foundation for systematically evaluating the AI cyber risk in realistic cyber offence operations.

AI安全漏洞挖掘模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。