构建评估多智能体AI蓝队能力的基准,推动更自主的网络安全运营中心。
Design Principles for the Construction of a Benchmark Evaluating Security Operation Capabilities of Multi-agent AI Systems
- 提出一套系统性设计原则,用于构建面向蓝队任务的多智能体评估基准。
- 设计涵盖五类大规模勒索攻击响应任务的SOC-bench概念框架。
- 填补现有研究空白,适合关注AI安全评估与自动化防御的研究者。
随着大语言模型和多智能体AI系统在网络安全运营中展现出日益增长的潜力,组织、政策制定者、模型提供商以及人工智能与网络安全领域的研究人员越来越希望量化此类AI系统的性能,以实现更自主的网络安全运营中心(SOC),减少人工干预。尽管近年来人工智能与网络安全社区已开发出若干评估多智能体AI红队能力的基准,但由于SOC中的操作主要由蓝队任务主导,若缺乏聚焦于蓝队操作的基准,就无法全面评估AI系统实现更自主SOC的能力。据我们所知,目前文献中尚无针对协同多任务蓝队AI系统的系统性基准。现有蓝队基准仅关注特定任务。本文旨在提出一套设计原则,用于构建名为SOC-bench的基准,以评估蓝队AI的能力。基于这些原则,我们提出了SOC-bench的概念设计,包含五类在大规模勒索软件攻击事件响应背景下的蓝队任务。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) and multi-agent AI systems are demonstrating increasing potential in cybersecurity operations, organizations, policymakers, model providers, and researchers in the AI and cybersecurity communities are interested in quantifying the capabilities of such AI systems to achieve more autonomous SOCs (security operation centers) and reduce manual effort. In particular, the AI and cybersecurity communities have recently developed several benchmarks for evaluating the red team capabilities of multi-agent AI systems. However, because the operations in SOCs are dominated by blue team operations, the capabilities of AI systems & agents to achieve more autonomous SOCs cannot be evaluated without a benchmark focused on blue team operations. To our best knowledge, no systematic benchmark for evaluating coordinated multi-task blue team AI has been proposed in the literature. Existing blue team benchmarks focus on a particular task. The goal of this work is to develop a set of design principles for the construction of a benchmark, which is denoted as SOC-bench, to evaluate the blue team capabilities of AI. Following these design principles, we have developed a conceptual design of SOC-bench, which consists of a family of five blue team tasks in the context of large-scale ransomware attack incident response.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。