用实时攻防赛构建新评测框架,解决大模型安全代理评估中的作弊问题。
CTFusion: A CTF-based Benchmark for LLM Agent Evaluation

- 基于实时攻防赛设计独立评测机制,防止多代理间数据泄露。
- 实测显示现有基准易被大模型利用搜索工具作弊,准确率下降超40%。
- 开源框架支持多种代理和赛事,适合安全与智能体研究者使用。
大型语言模型的进展推动了复杂多步骤任务的智能体系统发展,网络安全正成为其重要应用方向。当前主流评估方法依赖于捕获旗帜(CTF)竞赛基准,但这些基准复用已有题目,导致数据污染和潜在作弊风险。我们通过将网络搜索工具集成到现有智能体中,实证发现该问题严重:部分测试中模型准确率下降超过40%。为此,我们提出CTFusion——一种基于实时攻防赛的流式评估框架。该框架在单个团队账号下保持各智能体独立性,并仅转发每个挑战首个正确旗帜以降低竞争干扰。我们将其部署为在广泛使用的CTFd平台上运行的模型上下文协议(MCP)服务器,具备对不同赛事与智能体类型的通用适配性。在三个大模型、两个智能体和五个实时攻防赛上的实验表明,现有基准难以可靠评估大模型代理性能,而CTFusion能提供稳定可靠的评测方案。我们已开源CTFusion,以促进该领域未来研究。
原文摘要 · Abstract (English)
Recent advances in Large Language Models (LLMs) have enabled agentic systems for complex, multi-step tasks; cybersecurity is emerging as a prominent application. To evaluate such agents, researchers widely adopt Capture The Flag (CTF) benchmarks. However, current CTF benchmarks reuse existing challenges, which exposes them to data contamination and potential cheating. Notably, we confirmed these issues in practice by integrating web search tools into an existing agent. To address these limitations, we present CTFusion, a streaming evaluation framework built on Live CTFs. To achieve this, CTFusion preserves per-agent independence under a single team account and reduces competition impact by forwarding only the first correct flag per challenge. Moreover, we implement CTFusion as a Model Context Protocol (MCP) server on the widely used CTFd platform, which offers broad applicability to diverse CTF events and agent types. Through experiments with three LLMs, two agents, and five Live CTFs, we demonstrate that existing CTF benchmarks can be unreliable in assessing LLM-based agents, while CTFusion can serve as a robust solution for evaluating cybersecurity agents. We release CTFusion as open source to foster future research in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。