九类攻击链符号表示未必比五类更优,关键在因果状态区分。
Symbolic Attack Chain Generation from Atomic Red Team Techniques: An Empirical Study of Predicate Representation Granularity

- 用大模型转译+规划引擎推理,对比九类与五类符号粒度
- 16个技术测试中81.3%计划结果一致,仅1项因粒度提升获改进
- 高粒度主要提升解释结构,不直接影响攻击链可行性
自动化攻击链生成对现代网络安全至关重要,但人工构建难以随攻击行为扩展而扩展。尽管经典AI规划(如使用规划领域定义语言PDDL)提供了形式化自动化方法,其依赖于将技术准确转换为符号谓词。当前先进系统AURORA采用九类攻击动作链接模型(AALM),但该粒度是否必要尚未验证。本研究通过管道:大语言模型(LLM)进行转换,Fast Downward引擎执行确定性推理,对比全九类AALM与基于原子红队(ART)执行证据经验推导出的五类简化方案。由于九类域是五类域的重命名,两种方案在计划有效性与成本上保持一致;实质性测试聚焦于谓词类别解析能力。控制性A/B测试发现,粗粒度方案的计划虽通过所有有效性检查,却在操作层面错误:持有管理员权限并能用于网络登录是因果上不同的系统状态。16项技术测试中,81.3%计划结果一致,真正的谓词类别解析增益仅出现在1项技术中。结果表明,更高粒度主要增强计划论证的内部结构解析,而非攻击链本身的可行性。
原文摘要 · Abstract (English)
Automated attack chain generation is critical for modern cybersecurity, yet manual construction fails to scale as adversary behaviors expand. While classical AI planning using the Planning Domain Definition Language (PDDL) offers a formal method to automate this process, it relies on the accurate translation of techniques into symbolic predicates. Current state-of-the-art systems like AURORA employ a nine-category Attack Action Linking Model (AALM), but the necessity of this specific granularity remains unvalidated. This work investigates whether AURORA's nine-category taxonomy provides representational distinctions beyond those captured by a reduced, empirically derived scheme. Utilizing a pipeline where a Large Language Model (LLM) performs translation and the Fast Downward engine performs deterministic reasoning, the study compares the full nine-category AALM against a reduced five-category scheme derived empirically from Atomic Red Team (ART) execution evidence. Because the nine-category domain is constructed as a relabeling of the five-category domain, plan validity and cost are held identical between schemes by design; the substantive test of granularity's effect lies instead in the resulting predicate category resolution. There, a controlled A/B test isolates a case where a coarser scheme's plan passes every validity check while remaining operationally wrong: holding administrator privilege and being able to exercise it over a network logon prove to be causally distinct system states. Results from a sixteen-technique corpus show 81.3% identical plan outcomes across both schemes by construction, with a genuine predicate category resolution gain confined to a single technique out of sixteen. The findings suggest that higher granularity primarily enhances the internal structural resolution of a plan's justification rather than the viability of the generated attack chain itself.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。