arXiv:2508.00910cs.CRcs.CL2025-08被引 20

无需运行环境,用写题解模拟训练网络安全大模型。

Cyber-Zero: Training Cybersecurity Agents without Runtime

  • 用公开黑客比赛题解和角色模拟生成真实交互序列。
  • 在三个比赛基准上性能提升最高达13.1%。
  • 适合想低成本打造顶尖安全智能体的研究者。

大型语言模型(LLMs)在软件工程任务中表现优异,尤其在具备可执行运行环境时能有效解决GitHub问题。然而,在网络安全领域,挑战配置与执行环境常为临时或受限状态,难以获取。本文提出Cyber-Zero,首个无需运行环境的框架,通过公开的CTF写题解,结合人物驱动的LLM仿真,逆向推导运行行为,生成无需实际环境支持的高质量、长时序交互轨迹。利用这些合成轨迹训练的基于LLM的智能体,在InterCode-CTF、NYU CTF Bench和Cybench三个主流CTF基准上,性能相比基线模型最高提升13.1%。最佳模型Cyber-Zero-32B在开源权重模型中达到新SOTA,性能媲美DeepSeek-V3-0324与Claude-3.5-Sonnet等专有系统,且成本更低,证明了无运行环境轨迹合成可有效推动顶尖网络安全智能体的普惠化发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved remarkable success in software engineering tasks when trained with executable runtime environments, particularly in resolving GitHub issues. However, such runtime environments are often unavailable in other domains, especially cybersecurity, where challenge configurations and execution contexts are ephemeral or restricted. We present Cyber-Zero, the first runtime-free framework for synthesizing high-quality agent trajectories to train cybersecurity LLMs. Cyber-Zero leverages publicly available CTF writeups and employs persona-driven LLM simulation to reverse-engineer runtime behaviors and generate realistic, long-horizon interaction sequences without actual environments. Using trajectories synthesized by Cyber-Zero, we train LLM-based agents that achieve up to 13.1% absolute performance gains over baseline models on three prominent CTF benchmarks: InterCode-CTF, NYU CTF Bench, and Cybench. Our best model, Cyber-Zero-32B, establishes new state-of-the-art performance among open-weight models, matching the capabilities of proprietary systems like DeepSeek-V3-0324 and Claude-3.5-Sonnet while offering superior cost-effectiveness, and demonstrating that runtime-free trajectory synthesis can effectively democratize the development of state-of-the-art cybersecurity agents.

网络安全大模型轨迹生成CTF

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。