arXiv:2604.19354cs.AIcs.CR2026-04中稿 · AIWare'26 Benchmar…被引 2

评测大模型在真实攻防挑战中的表现,发现其能力仍有明显短板。

Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges

  • 构建开源基准DeepRed,在隔离环境中测试大模型攻防能力。
  • 采用分阶段计分法,最佳模型仅完成35%关键任务点。
  • 适合关注大模型安全应用潜力的研究者与开发者参考。

大语言模型代理被越来越多地提出用于自主网络安全任务,但在真实攻击场景中的能力仍不明确。我们提出了DeepRed,一个开源基准,用于在隔离虚拟化环境中评估基于大模型的代理在真实捕获旗帜(CTF)挑战中的表现。DeepRed将代理置于包含终端工具和可选网络搜索功能的Kali攻击环境,通过私有网络连接到目标挑战,并记录完整执行日志以供分析。为突破传统的成功/失败二元评价,我们引入基于公开题解提取的特定挑战检查点的半分制评分方法,并开发自动化‘摘要-判别’标注流程,从日志中判断检查点完成情况。利用DeepRed,我们在涵盖不同挑战类别的十项基于虚拟机的CTF挑战上,对十种商业可用的大模型进行了基准测试。结果表明,当前代理能力仍受限:最佳模型平均仅完成35%的检查点,对常见类型表现较好,而在需要非标准发现和长周期适应的任务上表现最弱。

原文摘要 · Abstract (English)

Large Language Model (LLM) agents are increasingly proposed for autonomous cybersecurity tasks, but their capabilities in realistic offensive settings remain poorly understood. We present DeepRed, an open-source benchmark for evaluating LLM-based agents on realistic Capture The Flag (CTF) challenges in isolated virtualized environments. DeepRed places an agent in a Kali attacker environment with terminal tools and optional web search, connected over a private network to a target challenge, and records full execution traces for analysis. To move beyond binary solved/unsolved outcomes, we introduce a partial-credit scoring method based on challenge-specific checkpoints derived from public writeups, together with an automated summarise-then-judge labelling pipeline for assigning checkpoint completion from logs. Using DeepRed, we benchmark ten commercially accessible LLMs on ten VM-based CTF challenges spanning different challenge categories. The results indicate that current agents remain limited: the best model achieves only 35% average checkpoint completion, performing strongest on common challenge types and weakest on tasks requiring non-standard discovery and longer-horizon adaptation.

大模型安全攻防测试评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。