arXiv:2606.28929cs.CRcs.LG2026-06

用网络安全检验生成式AI真本事,挑战远超语言与视觉任务。

Cybersecurity is the True Frontier for Generative AI Success or Failure

  • 以攻防对抗场景为测试场,整合数百种工具处理海量数据
  • 单个恶意样本达数十亿标记符,标注成本高且存在专家分歧
  • 强调低延迟、可解释性,适合追求落地可信AI的研究者

网络安全是多重机器学习问题的现实试验场,尤其在大型语言模型(LLMs)被用作自动化代理的背景下。其工作流程需协调数百种标准与定制化工具,数据规模极为庞大——单个恶意软件样本可视为数十亿标记符的序列。专家标注成本高昂且耗时,部分因对手(如资金充足的国家行为体)试图规避检测方法;即使资深专家也可能对标签存在分歧,导致真实标签模糊。模型部署后需每天处理数十亿条数据,低延迟对实际运营至关重要,且环境持续动态变化。此外,可解释性不可缺失:分析师需清晰推理依据以应对海量误报,并快速制定修复方案及追溯错误原因。综上,网络安全复杂度超过自然语言与计算机视觉,因此我们认为它是评估通用人工智能进展更优的试金石。

原文摘要 · Abstract (English)

Cybersecurity is a real-life test-bed for many machine learning problems at once, especially when considering modern strides in using Large Language Models (LLMs) to automate processes as ``agents.'' Cybersecurity workflows require orchestrating hundreds of standard and bespoke tools through various formats. The scale of cybersecurity data is enormous; for example, a single malware sample can be viewed as a sequence of billions of tokens. The cost of labeling any file by experts is enormous and labor-intensive, in part because an adversary (possibly a well-funded nation state actor) is attempting to subvert your detection methods. Even skilled experts may disagree on the correct label, creating ambiguity in what constitutes ground truth. When deployed, models must run quickly on billions of items a day, where low-latency is critical for operational success, in a continuously changing environment. In addition, explainability is not optional: analysts demand clear reasoning for model decisions to cope with the large number of false-positive alerts they face daily, and to quickly develop remediation and understand how something went wrong. In short, the amount of complexity cybersecurity is greater than that of natural language and computer vision, and thus we posit that cybersecurity is the better test-case for general AI progress than other, well-studied fields.

生成式AI网络安全可解释性大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。