评测大模型在网络安全任务中的表现,提供开源工具包。
DefenderBench: A Toolkit for Evaluating Language Agents in Cybersecurity Environments
- 构建可复现的网络安全评估环境,支持攻防与知识任务
- Claude-3.7-sonnet得分81.65领先,开放模型Llama 3.3 70B达71.81
- 模块化设计适配自定义模型和任务,适合安全与AI交叉研究者
大型语言模型(LLM)在语言理解与推理方面表现出色,但在网络安全领域的潜力尚未充分探索。我们提出DefenderBench,一个实用且开源的工具包,用于评估语言代理在攻击、防御及网络安全知识任务中的表现。该工具包涵盖网络入侵、恶意内容检测、代码漏洞分析与网络安全知识评估等环境,设计上兼顾低成本与易用性,确保评估公平严谨。我们使用标准化代理框架对多个前沿主流模型(包括开源与闭源模型)进行基准测试。结果显示,Claude-3.7-sonnet得分为81.65最高,其次为Claude-3.7-sonnet-think的78.40;最优开源模型Llama 3.3 70B得分为71.81,表现接近。DefenderBench采用模块化架构,支持灵活集成自定义模型与任务,促进结果可复现与公平比较。匿名版已公开于https://github.com/microsoft/DefenderBench。
原文摘要 · Abstract (English)
Large language model (LLM) agents have shown impressive capabilities in human language comprehension and reasoning, yet their potential in cybersecurity remains underexplored. We introduce DefenderBench, a practical, open-source toolkit for evaluating language agents across offense, defense, and cybersecurity knowledge-based tasks. DefenderBench includes environments for network intrusion, malicious content detection, code vulnerability analysis, and cybersecurity knowledge assessment. It is intentionally designed to be affordable and easily accessible for researchers while providing fair and rigorous assessment. We benchmark several state-of-the-art (SoTA) and popular LLMs, including both open- and closed-weight models, using a standardized agentic framework. Our results show that Claude-3.7-sonnet performs best with a DefenderBench score of 81.65, followed by Claude-3.7-sonnet-think with 78.40, while the best open-weight model, Llama 3.3 70B, is not far behind with a DefenderBench score of 71.81. DefenderBench's modular design allows seamless integration of custom LLMs and tasks, promoting reproducibility and fair comparisons. An anonymized version of DefenderBench is available at https://github.com/microsoft/DefenderBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。