arXiv:2602.02595cs.CRcs.AI2026-02被引 3

AI 代理将催生大规模精准攻击,防御者必须先学会‘黑客’才能自保。

To Defend Against Cyber Attacks, We Must Teach AI Agents to Hack

  • 用训练好的智能体自动发现真实网络中的漏洞,实现规模化攻击。
  • 现有防护手段对可控模型的攻击者无效,因他们可绕过安全限制。
  • 建议在受控环境中发展进攻性AI能力,用于提前防御真实威胁。

过去十年,网络安全依赖人力稀缺性,使攻击者只能针对高价值目标进行手动或通用自动化攻击。构建复杂漏洞利用需深厚专业知识和人工投入,导致防御方认为对手无法大规模定制攻击。然而,AI代理通过自动化跨数千目标的漏洞发现与利用,只需较低成功率即可盈利,打破这一平衡。当前开发者主要依赖数据过滤、安全对齐和输出防护来防止滥用,但这些措施对控制开源权重模型、绕过安全机制或独立开发攻击能力的对手无效。本文认为,基于AI代理的网络攻击不可避免,必须从根本上改变防御策略。我们指出现有防护无法应对自适应攻击者,并主张防御方必须发展进攻性安全情报。为此提出三项行动:构建覆盖攻击全生命周期的综合性基准;从工作流驱动转向训练有素的智能体,实现对真实环境漏洞的大规模发现;实施治理机制,将进攻型智能体限制在审计过的网络靶场中,按能力层级分阶段发布,并将成果提炼为仅具防御功能的智能体。强烈建议将进攻性AI能力视为必要防御基础设施,在对手掌握前于受控环境中先行掌握。

原文摘要 · Abstract (English)

For over a decade, cybersecurity has relied on human labor scarcity to limit attackers to high-value targets manually or generic automated attacks at scale. Building sophisticated exploits requires deep expertise and manual effort, leading defenders to assume adversaries cannot afford tailored attacks at scale. AI agents break this balance by automating vulnerability discovery and exploitation across thousands of targets, needing only small success rates to remain profitable. Current developers focus on preventing misuse through data filtering, safety alignment, and output guardrails. Such protections fail against adversaries who control open-weight models, bypass safety controls, or develop offensive capabilities independently. We argue that AI-agent-driven cyber attacks are inevitable, requiring a fundamental shift in defensive strategy. In this position paper, we identify why existing defenses cannot stop adaptive adversaries and demonstrate that defenders must develop offensive security intelligence. We propose three actions for building frontier offensive AI capabilities responsibly. First, construct comprehensive benchmarks covering the full attack lifecycle. Second, advance from workflow-based to trained agents for discovering in-wild vulnerabilities at scale. Third, implement governance restricting offensive agents to audited cyber ranges, staging release by capability tier, and distilling findings into safe defensive-only agents. We strongly recommend treating offensive AI capabilities as essential defensive infrastructure, as containing cybersecurity risks requires mastering them in controlled settings before adversaries do.

AI攻防网络安全智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。