arXiv:2607.25425cs.AIcs.CR2026-07

大模型正颠覆网络安全竞赛,论文提出公平竞赛的四维防护框架。

The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play

论文配图:The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play
图 1 · 摘自论文原文
  • 通过多方法研究,界定大模型在密码学、网络与二进制攻防中的自动化边界。
  • 发现易到中等难度题已可被大模型稳定破解,但部分子领域仍具挑战性。
  • 为不同目标的竞赛设计了分层防护方案,适合组织者与安全教育者参考。

Capture the Flag(CTF)竞赛是网络安全领域最有效的实践训练场,涵盖密码学、网络攻击与二进制漏洞利用等技能培养。如今大语言模型(LLMs)可在极少人工干预下解决越来越多的题目,引发对公平性、排名有效性及参赛学习价值的深刻质疑。本文基于混合方法研究,整合公开基准测试(包括一次政府评估)、三类挑战的实战案例、公共讨论区的结构化观察,以及对资深选手与组织者的半结构化访谈,绘制出当前人机能力边界。结果表明,密码学、网络与二进制攻防中的简单与中等难度题目已可被可靠自动化;而部分细分领域仍具抗大模型能力。社区对是否允许使用AI的分歧,根源于对竞赛本质的未明确认知。为此,本文提出包含四要素的保障框架:分级竞赛制度、抗大模型挑战设计、用于调查的遥测数据、草案社区行为准则,并配套一个决策工具,将防护组合与竞赛目标相匹配。该结论亦适用于所有以成果证明能力的网络安全场景。

原文摘要 · Abstract (English)

Capture the Flag (CTF) competitions are among cybersecurity's most effective training grounds, developing practical skill across cryptography, web exploitation, and binary exploitation. Large language models (LLMs) can now solve a growing share of challenges with minimal human input, raising urgent questions about fairness, the validity of rankings, and whether participation still delivers the learning that justifies the effort. This paper reports a mixed-methods study of LLM impact on modern CTFs, combining a synthesis of published benchmarks, including a recent government evaluation, case studies of live competition across three challenge categories, structured observation of the public channels where the community debates AI use, and semi-structured interviews with experienced players and organisers. We map the current human-machine capability boundary by category, showing that easy and intermediate challenges in cryptography, web, and binary exploitation are now reliably automated while narrower sub-categories continue to resist. We find that community disagreement about whether AI should be permitted is downstream of an undeclared prior question: what a competition is for. Against this backdrop we contribute a four-component safeguard framework, combining tiered competition divisions, LLM-resistant challenge design, telemetry used investigatively, and a draft community code of conduct, together with a decision tool that ties the combination of safeguards to a competition's declared purpose. The argument reaches beyond CTFs to any setting in cybersecurity where a demonstrated result is taken as evidence of an underlying ability.

网络安全大模型竞赛公平智能攻防

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。