arXiv:2509.14289cs.AIcs.CL2025-09EMNLP被引 2

测试大模型在渗透攻击中的真实表现,发现关键能力可显著提升攻防效率。

From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration Testing

  • 通过五种能力增强,系统评估大模型在渗透测试中的表现
  • 模块化设计配合增强后,复杂任务成功率明显提升
  • 适合安全研究者和红队人员参考实战优化策略

大型语言模型(LLMs)被越来越多用于自动化或辅助渗透测试,但其在不同攻击阶段的有效性和可靠性尚不明确。本文对多种基于LLM的智能体(从单智能体到模块化设计)进行了全面评估,涵盖真实渗透测试场景,测量实际性能并识别重复性失败模式。通过针对性增强五个核心功能:全局上下文记忆(GCM)、智能体间通信(IAM)、上下文驱动调用(CCI)、自适应规划(AP)和实时监控(RTM),分别实现上下文一致性、组件协同、工具使用精准性、多步策略规划与容错、动态响应。结果表明,尽管部分架构天然具备部分特性,但针对性增强能显著提升模块化智能体在复杂、多步骤及实时渗透任务中的表现。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to automate or augment penetration testing, but their effectiveness and reliability across attack phases remain unclear. We present a comprehensive evaluation of multiple LLM-based agents, from single-agent to modular designs, across realistic penetration testing scenarios, measuring empirical performance and recurring failure patterns. We also isolate the impact of five core functional capabilities via targeted augmentations: Global Context Memory (GCM), Inter-Agent Messaging (IAM), Context-Conditioned Invocation (CCI), Adaptive Planning (AP), and Real-Time Monitoring (RTM). These interventions support, respectively: (i) context coherence and retention, (ii) inter-component coordination and state management, (iii) tool use accuracy and selective execution, (iv) multi-step strategic planning, error detection, and recovery, and (v) real-time dynamic responsiveness. Our results show that while some architectures natively exhibit subsets of these properties, targeted augmentations substantially improve modular agent performance, especially in complex, multi-step, and real-time penetration testing tasks.

大模型安全渗透测试智能体系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。