arXiv:2603.10041cs.CRcs.LG2026-03

测试自适应攻击代理在未知网络地址分配下的泛化能力,发现仅元学习能有效应对变化。

Evaluating Generalization Mechanisms in Autonomous Cyber Attack Agents

  • 在固定企业环境中改变主机IP段,测试攻击代理的泛化能力
  • 基于提示的预训练LLM代理在未见配置下成功率最高
  • 尽管性能好,但存在重复动作、无效循环等实际问题

自主进攻型智能体通常无法在训练网络之外的环境中迁移。我们隔离了一个最小但根本性的变化——在原本固定的企事业场景中出现未见过的主机/子网IP重新分配,并在NetSecGame环境中评估攻击者泛化能力。代理在五个IP范围变体上训练,在第六个未见过的变体上测试;只有元学习代理可在测试时自适应。比较了三类代理(传统强化学习、自适应代理、基于LLM的代理),并使用基于动作分布的行为分析与可解释AI方法定位失败模式。部分自适应方法虽有部分迁移能力,但在未见重分配下仍出现显著性能下降,表明即使地址空间的变化也会破坏长时程攻击策略。在本评估协议及代理假设下,提示驱动的预训练LLM代理在保留的重分配任务中达到最高成功率,但代价是增加推理计算开销、降低透明度,以及出现重复或无效动作循环等实际失效模式。

原文摘要 · Abstract (English)

Autonomous offensive agents often fail to transfer beyond the networks on which they are trained. We isolate a minimal but fundamental shift -- unseen host/subnet IP reassignment in an otherwise fixed enterprise scenario -- and evaluate attacker generalization in the NetSecGame environment. Agents are trained on five IP-range variants and tested on a sixth unseen variant; only the meta-learning agent may adapt at test time. We compare three agent families (traditional RL, adaptation agents, and LLM-based agents) and use action-distribution-based behavioral/XAI analyses to localize failure modes. Some adaptation methods show partial transfer but significant degradation under unseen reassignment, indicating that even address-space changes can break long-horizon attack policies. Under our evaluation protocol and agent-specific assumptions, prompt-driven pretrained LLM agents achieve the highest success on the held-out reassignment, but at the cost of increased inference-time compute, reduced transparency, and practical failure modes such as repetition/invalid-action loops.

网络安全攻击模拟泛化能力LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。