arXiv:2606.08168cs.CRcs.AI2026-06

首个评估自主防御代理在商业EDR中表现的框架,揭示仿真与现实差距。

Closing the Sim-to-Real Gap: An Evaluation Framework for Autonomous Cyber Defense Configuration of Commercial EDR

  • 构建基于GOAD实验室的评估框架,用AI攻防测试商业EDR
  • LLM驱动的防御代理在真实环境中表现受EDR自治行为影响
  • 发现商业EDR日志不适用于科学评测,需区分策略归属

主流商业终端检测与响应(EDR)产品已从人工配置规则转向由自主AI组件协同甚至替代人工策略。如今,利用商业EDR作为加固工具的自主防御代理,面对的不再是静态工具,而是一个黑箱式、可自主决策的系统。本文提出首个针对此类自主防御代理在商业EDR中表现的评估框架。我们在Game of Active Directory(GOAD)实验环境中,以Horizon3.ai的NodeZero为自主渗透测试器,微软Defender XDR为EDR平台,对两个大型语言模型(Claude Sonnet 4.6和Cisco Foundation-Sec-8B)驱动的防御代理进行基准测试。研究发现三个现有仿真与开源EDR评估无法揭示的问题:(i) 商业EDR遥测数据专为安全运营中心(SOC)分析师工作流设计,而非科学评测;(ii) 必须对每条策略进行溯源,以区分防御代理行为与自主EDR自身动作;(iii) EDR的自主行为在评估周期内存在动态变化。这些发现凸显企业级防御中的仿真到现实差距,并推动针对黑箱自主工具的评估方法发展。

原文摘要 · Abstract (English)

Leading commercial endpoint detection and response (EDR) products have shifted from operator-configured rule sets to multi-component systems where autonomous AI components operate alongside, and increasingly in place of, operator-deployed policies. Autonomous defense agents using commercial EDR as their hardening tool are no longer tuning a passive tool, but a black-box autonomous system capable of making vendor-specific decisions. We present the first evaluation framework for autonomous defense agents hardening commercial EDR. We instantiate it in a Game of Active Directory (GOAD) lab with Horizon3.ai's NodeZero as the autonomous pentester and Microsoft Defender XDR as the EDR. We run a sample benchmark of defense agents with two large language model (LLM) backbones (Claude Sonnet 4.6 and Cisco Foundation-Sec-8B). We report three lessons learned that neither simulation nor open-source-EDR evaluation can surface: (i) commercial EDR telemetry is engineered for Security Operations Center (SOC) analyst workflows rather than scientific benchmarking; (ii) the importance of per-policy attribution to separate defense agent actions from autonomous EDR actions; and (iii) the EDR's autonomous behavior varies during the evaluation window. Together, these findings highlight a sim-to-real gap for enterprise defense and motivate evaluation methodology for benchmarking autonomous defense agents in environments with black-box, autonomous tools.

自主防御EDR评估攻防对抗AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。