arXiv:2603.02297cs.CRcs.AI2026-03中稿 · ICLR被引 6

评测大模型在未知漏洞上的自动修复能力,发现当前顶级模型仍难胜任。

ZeroDayBench: Evaluating LLM Agents on Unseen Zero-Day Vulnerabilities for Cyberdefense

  • 构建新基准ZeroDayBench,测试模型发现并修补22个真实开源漏洞
  • 三款前沿模型在任务中表现有限,平均修复成功率不足50%
  • 揭示模型在主动防御中的行为缺陷,为安全增强提供方向

大型语言模型(LLMs)正被越来越多地用作自主参与代码库的软件工程代理。其主要优势在于能够发现并修补所监管代码库中的安全漏洞。为评估此类代理在该领域的实际能力,我们提出ZeroDayBench基准,要求LLM代理在22个新开源代码库中发现并修补22个新型关键漏洞。研究聚焦于三款主流前沿智能体模型:GPT-5.2、Claude Sonnet 4.5和Grok 4.1。结果表明,当前前沿模型尚无法自主完成上述任务,且观察到若干行为模式,提示了在主动网络安全防御领域提升模型能力的关键路径。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly being deployed as software engineering agents that autonomously contribute to repositories. A major benefit these agents present is their ability to find and patch security vulnerabilities in the codebases they oversee. To estimate the capability of agents in this domain, we introduce ZeroDayBench, a benchmark where LLM agents find and patch 22 novel critical vulnerabilities in open-source codebases. We focus our efforts on three popular frontier agentic LLMs: GPT-5.2, Claude Sonnet 4.5, and Grok 4.1. We find that frontier LLMs are not yet capable of autonomously solving our tasks and observe some behavioral patterns that suggest how these models can be improved in the domain of proactive cyberdefense.

安全检测大模型代理漏洞修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。